GLM-5.3 and GLM-5.3 Flash Are Now on Multos — Here's What Makes Them Special
Z.ai's latest models bring 1M context, native vision, and coding scores that rival models costing ten times more. We added them across four providers so you always have a fallback.
Two new models from Z.ai landed on Multos this week: GLM-5.3 on Baseten, and GLM-5.3 Flash on Together AI — joining the existing Novita, Baseten, and DeepInfra routes. If you've been sleeping on the GLM family, this is a good moment to pay attention.
What Z.ai Actually Built Here
GLM-5.3 dropped on August 14, 2026. It runs on the same 744-billion-parameter base as GLM-5.2 — no new architecture, no retraining from scratch. Every gain you see on benchmarks came purely from what Z.ai calls environment scaling: running the model through a much wider and harder set of task environments during post-training. On coding, it scored 28.3 on Terminal Bench 3.0 — open-source state of the art, up from 4.6 on the previous version. On FrontierSWE it went from 67.5 to 78.1. DeepSWE jumped from 46.2 to 66.9. The vulnerability discovery benchmark (CyberGym) more than doubled versus GLM-5.2. These are not incremental improvements. The full model is API-only for now — weights are gated.
GLM-5.3 Flash: Open Weights, Native Multimodal, One-Tenth the Price
GLM-5.3 Flash released on August 26, 2026. It's a genuinely different model: 320 billion total parameters, 18 billion active per token, built on a newly trained base with a hybrid sparse+linear attention architecture that cuts long-context serving costs significantly. It's the first natively multimodal model in the GLM-5 series — processing text, images, and video out of the box, not as an add-on. The context window sits at 1 million tokens. Weights are MIT licensed, so you can use it commercially, fine-tune it, and ship it with no restrictions. DeepSWE score of 63.4. HLE (frontier reasoning) at 55.3. One-tenth the price of GLM-5.3 while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Why We Added Together AI as a Fourth Provider
Multos already had GLM-5.3 Flash on Novita, Baseten, and DeepInfra. Adding Together AI gives the auto-router a fourth fallback path — so if any single provider is saturated or having an outage at peak hours, your agent keeps running. All four providers serve the same model ID (zai-org/GLM-5.3-Flash), so you get consistent output regardless of which one handles a given request. GLM-5.3 Flash sits in the Lite tier on Multos — same cost band as DeepSeek V4 Flash and MiMo V2.5. GLM-5.3 (the full flagship) is on Baseten in the Starter tier.
What You Can Do With It on Multos
Large codebase refactoring: load your entire project into context — up to 1M tokens — and ask it to refactor a pattern across every file. The long-context architecture means it doesn't lose track of what it changed twelve files ago. Vision and code combined: paste a screenshot of a UI bug, a Figma mockup, or a diagram alongside your code and ask GLM-5.3 Flash to fix or implement it. Multi-step agentic runs: the 63.4 DeepSWE score tells you this model holds a plan together across many sequential steps. Document analysis at scale: a 400-page technical spec, a stack of PDFs, a long API reference — all in one pass, no chunking required. Cybersecurity analysis: Z.ai specifically trained on vulnerability discovery environments, so this model understands security domains better than most alternatives at this price point.
How to Use It
Open the model dropdown on Multos and look under Together AI for GLM-5.3 Flash (new), or under Novita AI, Baseten, or DeepInfra for the existing routes. GLM-5.3 (the full model) is under Baseten. All four GLM-5.3 Flash providers are on the same Lite tier — pick whichever, or leave it on Auto and the router handles the rest. The model supports a reasoning_effort parameter with three levels: low, high, and max. It defaults to max, which is what Multos uses. For faster responses where you don't need deep reasoning, low cuts latency noticeably. If you've been defaulting to DeepSeek V4 Flash for quick agentic tasks, GLM-5.3 Flash is worth trying as a direct swap — same tier, same price band, but with native vision support and a 1M context window.
Ready to Get Started?
GLM-5.3 Flash is live on Multos Lite and above via Novita, Baseten, DeepInfra, and Together AI. Native image and video input, 1M token context, frontier-tier intelligence at flash-tier pricing. GLM-5.3 (full flagship) is on Baseten at the Starter tier. Select from the model dropdown or let Auto Router pick — available from the $20/month Lite plan upward.
Start Building FreeRelated Articles
GLM-5.3 Flash: The Model That Was Hiding in Plain Sight
For six days, one of the most capable models on OpenCode and OpenRouter had no name. It was just called Ox Alpha. Then Z.ai revealed it: GLM-5.3 Flash, a 320B MoE model with a 1M token context, native multimodal input, MIT-licensed weights — and a price tag that makes you read it twice.
Read more →Qwen 3.8 Flash Is Now on Multos — A 125B Model That Costs Less Than DeepSeek V4 Flash
Alibaba's newest model activates just 6B parameters per token, beats DeepSeek V4 Pro on SWE-bench Pro at 62.5 vs 55.4, and costs $0.15 per million tokens. We added it to Novita AI at the Lite tier.
Read more →