Z.ai
Important
Z.aiOpen SourceModelsChinaMultimodalZ.ai Releases GLM-5.3-Flash: Open-Weight Multimodal Model Running on Chinese Chips
August 26, 20263 min read
Z.ai has launched GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series. The 320B-total / 18B-active parameter model features a 1M-token context window, is released under the MIT License, and was previously previewed as Ox Alpha. It runs entirely on Chinese AI chips and is positioned as a high-performance, low-cost open-weight option.
Why it matters
The release strengthens the open-weight competitive landscape and demonstrates that competitive multimodal and coding performance can be delivered efficiently on non-Nvidia Chinese silicon, with fully open weights.
Z.ai (the international brand of Chinese lab Zhipu) has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. The model uses a hybrid linear and sparse attention architecture with 320 billion total parameters and 18 billion activated, a 1-million-token context window, and is released under the MIT License with open weights.
Previously previewed anonymously as Ox Alpha on OpenRouter, GLM-5.3-Flash is designed for coding, agentic workflows, office documents, and financial research. Z.ai emphasizes that the model runs entirely on Chinese AI chips and delivers strong performance at significantly lower cost than many frontier alternatives.
API pricing is highly competitive (approximately $0.15 input / $0.50 output per million tokens). The combination of open weights, native multimodality, long context, and non-Nvidia silicon makes this one of the more strategically notable open releases of the month.