# ZhipuAI Releases GLM-5.3-Flash Open-Weight Model at 1/40th Claude Opus Cost

> ZhipuAI launches GLM-5.3-Flash as the first multimodal GLM-5 model with hybrid attention architecture, 1M context, and MIT-licensed weights, matching top-tier intelligence while running on domestic Chinese chips at sharply lower prices.

*Published 2026-08-27 · By Marcus Vance*

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series from ZhipuAI, featuring 320B total parameters with 18B active parameters, a 1M-token context window, and full weights released under the MIT license on Hugging Face.

The release of GLM-5.3-Flash by ZhipuAI introduces a new era of accessible high-performance AI models that were previously only available through closed APIs from major Western companies. This model stands out as the initial natively multimodal offering in the GLM-5 lineup, allowing users to process text, images, and other data types seamlessly within the same framework. It combines advanced capabilities with cost efficiency that challenges existing market leaders by offering similar intelligence levels at a dramatically reduced price point, making it attractive for both individual developers and large enterprises seeking to integrate AI into their workflows without incurring prohibitive expenses. The full public release of weights enables independent verification and customization that closed models cannot match.

## Background and Context of the GLM-5 Series

ZhipuAI has been developing the GLM series over several iterations with a focus on balancing scale, efficiency, and openness. Previous versions like GLM-5.2 provided strong performance on coding and reasoning tasks but lacked native multimodal support from the initial training phase. The series has emphasized expanding context windows and improving hardware efficiency to meet growing demands for complex, long-context applications in research and industry. The transition to GLM-5.3 marks a deliberate shift toward multimodal integration from the ground up rather than retrofitting capabilities after pre-training.

The development of GLM-5 models has emphasized open-source principles to accelerate collective progress in the field. This approach allows researchers and developers worldwide to build upon the work, fine-tune for specific domains, and contribute improvements back to the community. Earlier releases established ZhipuAI as a key player in the Chinese AI ecosystem, and the current model extends that trajectory by combining frontier-level intelligence with unprecedented cost reductions and hardware flexibility.

## Introduction of GLM-5.3-Flash and Its Origins as Ox Alpha

GLM-5.3-Flash was initially tested anonymously under the name Ox Alpha on platforms such as OpenRouter and OpenCode. During this period, it became the most popular model of the week with all traffic served exclusively on Chinese AI chips. The model recorded over 60 trillion tokens in usage, demonstrating its robustness and appeal to users across diverse workloads including coding, agentic tasks, and multimodal processing. This anonymous testing phase allowed ZhipuAI to gather extensive real-world performance data without preconceptions about its origins.

The strategy of anonymous deployment proved effective as the model gained rapid traction among developers seeking high-performance options at lower costs. Feedback from the testing period informed final optimizations before the official launch. The full reveal now provides the community with complete access to the weights, documentation, and deployment guides, removing previous barriers to adoption and experimentation.

## Technical Specifications and Hybrid Architecture

GLM-5.3-Flash features 320 billion total parameters with only 18 billion active parameters per token. This sparse activation design contributes to its efficiency while maintaining high capability across modalities. It is the first open-source frontier model to use a hybrid architecture combining sparse attention and linear attention, a choice that directly addresses the computational bottlenecks of traditional transformer designs. The architecture was trained on a 30 trillion token multimodal pre-training corpus that equips the model for integrated text and vision tasks.

The hybrid approach reduces attention computation by 3.01 times and KV cache size by 4.44 times compared to GLM-5.3. These efficiency gains enable longer context handling without proportional increases in memory or compute requirements. The model supports a context window of 1,048,576 tokens, allowing it to process entire codebases, lengthy documents, or extended multimodal sequences in a single pass. Full weights are publicly available under the MIT license, facilitating both commercial and research use cases.

Performance comparison on key benchmarks showing consistent improvements of GLM-5.3-Flash over GLM-5.2BenchmarkGLM-5.3-FlashGLM-5.2DeepSWE v1.163.446.2AutomationBench v1.0.648.826.2

## Benchmark Results and Intelligence Index Score

On the Artificial Analysis Intelligence Index v4.1.1, GLM-5.3-Flash scored 57. This places it on par with Claude Opus 4.8 while delivering the result at substantially lower inference costs. The model also shows clear gains over GLM-5.2 in specialized benchmarks that measure software engineering and automation capabilities. These results validate the effectiveness of the hybrid architecture and the multimodal pre-training approach.

## Infrastructure Deployment on Domestic Chips

The inference service for GLM-5.3-Flash operated on a cluster of more than 100,000 domestic chips. This setup achieved hardware efficiency comparable to mainstream Nvidia GPUs while demonstrating the viability of non-Western silicon for frontier-model workloads. The company reported a three times improvement in end-to-end serving performance relative to prior configurations. This deployment marks a practical milestone in scaling open models on domestic infrastructure.

> This proves that domestic chips can fully and efficiently support frontier-model inference in large-scale scenarios.Zhipu AI / Z.ai, Company statement

## Pricing and Cost Advantages

The pricing is set at $0.15 per million input tokens and $0.50 per million output tokens. This represents one-tenth the price of GLM-5.3 and roughly one-fortieth the cost of Claude Opus 4.8. Such pricing makes frontier capabilities more accessible to a broader range of users and organizations that previously could not justify the expense of closed frontier models. The combination of open weights and low per-token costs creates new opportunities for experimentation at scale.

## Market and Stakeholder Implications

The open release under MIT license on Hugging Face enables widespread adoption and fine-tuning by researchers, startups, and enterprises. Stakeholders in the AI industry can now experiment with a high-performing model without high licensing fees or usage restrictions that accompany closed systems. This development may accelerate innovation in agentic and coding applications where long context and multimodal understanding provide competitive advantages.

- Democratization of access to frontier models through open weights and low pricing.
- Validation of domestic chip infrastructure for large-scale AI inference at frontier scale.
- Potential for reduced dependency on foreign hardware and software ecosystems in AI deployment.
- Encouragement of further competition in the multimodal AI space through replicable architecture choices.

## Expert Reactions and Company Statements

Company announcements highlight the model's design for low cost while maintaining high performance across benchmarks and real-world workloads. The statement emphasizes the multimodal pre-training corpus of 30 trillion tokens that underpins the model's cross-modal capabilities. This corpus supports the model's performance in coding, agentic benchmarks, and multimodal tasks at one-tenth the price of prior GLM-5 releases.

## What's Next for the GLM-5 Series

With the release of GLM-5.3-Flash, ZhipuAI is positioned to continue expanding the series with additional multimodal variants and further efficiency optimizations. Future models may build on the hybrid architecture to further reduce compute requirements while preserving or improving intelligence scores. The success of this open-weight approach could influence other companies to follow suit in releasing frontier models under permissive licenses.

The availability of full weights allows the community to contribute to improvements, create specialized fine-tunes, and explore novel applications. Researchers can investigate long-context reasoning and multimodal integration using the one-million-token window. Overall, this release signals a shift toward more inclusive AI development that prioritizes both performance and accessibility.

## Sources

1. [GLM-5.3-Flash scored 57 on the Artificial Analysis Intelligence Index v4.1.1 at a discounted cost of $0.045 per task.](https://www.zhipuai.cn/zh/research/163)
2. [We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.](https://huggingface.co/zai-org/GLM-5.3-Flash)
3. [GLM-5.3-Flash has 320B total parameters, with 18B activated parameters. It is the first open-source frontier model to adopt a hybrid architecture combining sparse attention and linear attention. Compared with GLM-5.3, it reduces attention computation and KV cache size by 3.01× and 4.44×, respectively. Outperforms GLM-5.2 on DeepSWE v1.1 (63.4 vs. 46.2) and AutomationBench (48.8 vs. 26.2).](https://docs.z.ai/guides/vlm/glm-5.3-flash.md)
4. [The company said that GLM-5.3-Flash has been officially released as an open-source model under the MIT license. It supports a 1-million-token context window. Inference service ran on a cluster of more than 100,000 domestic chips with hardware efficiency comparable to mainstream Nvidia GPUs and 3x end-to-end serving performance improvement. This proves that domestic chips can fully and efficiently support frontier-model inference in large-scale scenarios.](https://www.globaltimes.cn/page/202608/1369157.shtml)

---
Source: https://aiintelreport.com/frontier-models/zhipuai-glm-5-3-flash-open-weight-release
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
