The latest in AI, every dayAI News
← AI News

September 2, 2026 · Alibaba / CellCog

Qwen3.8-Max-0902 Improves All Coding Benchmarks at the Same Price

My take: Alibaba released Qwen3.8-Max-0902 today, an improved snapshot of its most capable model, with notable gains in coding and agentic benchmarks. TerminalBench 3.0 jumped from 11.3 to 29.0, DeepSWE 1.1 from 56.6 to 69.3, and QwenSWEbench V2 from 55.1 to 70.0. Pricing remains the same: $2 per million input tokens and $6 per million output.

Worth noting: several of these benchmarks were designed by Alibaba's own team, so as both judge and party, waiting for independent evaluations before treating those numbers as definitive is the right call. What is confirmed without needing external validation is the price: same as before, with a model that gets better.

If you use Qwen in your agent stack or coding workflows, do you have a plan to test this new snapshot against your actual use case before migrating?

Read at the source: Alibaba / CellCog ↗

Want to use these tools? See the unbiased reviews or back to the news.