← AI News

August 15, 2026 · SiliconANGLE

Z.ai Ships GLM-5.3 Without Retraining the Base Model: 50% Better at Coding and Leading Cybersecurity Benchmarks

My take: Z.ai released GLM-5.3 yesterday without retraining the base model: it started from the same 743-billion-parameter checkpoint used for GLM-5.2 and achieved all the improvement in post-training. On Terminal-Bench 3.0, the long-horizon coding score jumped from 4.6 to 28.3. On CyberGym, the third-party cybersecurity benchmark, it reached 84.5%. The model also outperforms Kimi K3 on Humanity's Last Exam with tools (62.5% vs. 56%).

What gives this story technical significance is the lesson it leaves: there is still enormous room to improve existing models without retraining from scratch. Z.ai admitted that the cybersecurity jump exceeded their own expectations, meaning they did not fully plan it — it emerged from the process. That kind of surprise result in security capabilities deserves attention.

For teams working on long-horizon code automation or systems defense, GLM-5.3 has benchmarks worth evaluating. Open weights arrive in approximately two weeks, so by the end of August you can run it locally. The 1-million-token context window and 40 billion active parameters per token make it viable for long tasks without a prohibitive per-inference cost.

Do you already have GLM-5.3 on your list of models to evaluate for long-horizon coding or defensive cybersecurity tasks?

Read at the source: SiliconANGLE ↗

Want to use these tools? See the unbiased reviews or back to the news.