Z.ai Reveals GLM-5.3-Flash: 320B MoE Model Approaching Claude Opus on Benchmarks
Z.ai revealed that the anonymously tested Ox Alpha model is GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model with only 18 billion active parameters. The multimodal model approached Claude Opus 4.8 on coding and agentic benchmarks while designed for extreme efficiency and ultra-low-cost inference. All inference traffic was served on Chinese AI chips, demonstrating viability of alternative hardware stacks. The model went viral after being released on OpenCode and OpenRouter.
Why it matters
💻 Developer · This model proves you don't need 10B active parameters for strong coding performance. The sparse MoE architecture is open source—study how it routes computation and consider this approach for your inference infrastructure.
📦 Product · Ultra-efficient inference at scale changes unit economics for AI features. If GLM-5.3-Flash really costs 95% less than alternatives, you can offer features at price points previously impossible.
🎨 Design · Multimodal inference on cheaper chips means you can build visual features into more products. Test whether this model's speed lets you build real-time interactive experiences.
📈 Business · This breaks the AI compute oligopoly narrative. If models work on non-NVIDIA hardware, you have vendor negotiating power and geopolitical optionality. The efficiency story also matters for sustainable margin.
🤔 Just Curious · This is the first major multimodal MoE that actually works better than dense models at smaller parameter counts. The mystery release and viral adoption reveals how good models now spread through decentralized channels, not corporate announcements.
Sources: Ox-Alpha Revealed as GLM-5.3-Flash, What Z.ai's Ox Alpha reveals about AI economics