The latest in AI, every dayAI News
← AI News

September 5, 2026 · Z.ai

Z.ai Launches GLM-5.3-Flash: Open 320B Multimodal Model With 1M-Token Context

My take: Z.ai published GLM-5.3-Flash, the first natively multimodal model in the GLM-5 family: 320 billion parameters in a MoE architecture that activates only 18 billion at a time, a one-million-token context window, support for text, image, and video input, and open weights under the MIT license.

For those building projects or applying AI at work, a model of this caliber at $0.15 per million input tokens (with a launch price of $0.075 through September 9) is a real option for workflows that previously cost ten times more. Combined with the long context and multimodality, it is useful for analyzing long documents, processing video, or running agents that handle large volumes of information repeatedly.

What I find most relevant about this launch is not just the price: it is that frontier open-source models are now multimodal, affordable, and downloadable. That changes the calculus for any project where API cost or vendor dependency was the bottleneck.

If you have a workflow you currently limit because of API cost or vendor lock-in, how would you approach it with a model like this running on your own server?

Read at the source: Z.ai ↗

Want to use these tools? See the unbiased reviews or back to the news.