Frontier Models
Tencent Hunyuan Hy4 Preview Launches 770B MoE Frontier Model
The August 28 2026 release delivers competitive engineering task performance at lower per token costs while integrating across productivity applications and third party platforms.
Hunyuan Hy4 preview is a next generation mixture of experts large language model developed by Tencent with 770 billion total parameters.
Tencent released and open sourced the Hunyuan Hy4 preview on August 28 2026. The release includes the model weights and makes it accessible through various platforms. This action positions the company in the competitive landscape of large language models. The preview version focuses on practical applications in professional settings.
Availability includes API access on Tencent Cloud TokenHub. It is also listed on OpenRouter for broader reach. Users can interact with the model through integrated tools such as WorkBuddy and CodeBuddy. Additional platforms include Yuanbao and ima for specialized uses.
Tencent has been developing the Hunyuan series for several years. The series aims to provide powerful tools for both consumer and enterprise applications. The Hy4 preview represents the latest advancement in this line.
Previous models like Hy3 have received positive reception in the market. The extension of free access to Hy3 shows commitment to user support during transitions. The new model builds on lessons from earlier iterations.
The focus on internal productivity benchmarks reflects the company's emphasis on practical utility. Models are tested on tasks relevant to real business operations. This approach differs from purely academic evaluations.
What are the technical specifications of the Hunyuan Hy4 preview model?
The Hunyuan Hy4 preview operates with a total of 770 billion parameters. Of these 49 billion activate for each token during processing. The architecture employs a mixture of experts approach to optimize efficiency. A context window exceeding 1 million tokens permits handling of very long sequences of data. This setup is particularly suited for tasks requiring extensive context retention.
The mixture of experts structure means that the model routes tokens to specialized sub networks. This selective activation reduces the computational load while preserving model capacity. It represents a common approach in scaling large models efficiently.
The 1 million token context allows the model to maintain context over very long documents or code bases. This is critical for tasks like analyzing entire code repositories or comprehensive reports. Users in research fields can input large datasets for analysis.
Development relied on high quality data co built with Tencent experts. The domains covered software engineering gaming finance and security. This targeted data collection enhances performance in those areas. The model was trained to handle domain specific challenges effectively.
An notable aspect is the autonomous optimization of its own inference system. This self improvement led to a 31.8 percent increase in end to end throughput. The baseline comparison shows significant gains in operational efficiency. Such optimizations reduce the resources needed for deployment.
The total parameter count of 770 billion places it among the larger models currently available. However the sparse activation keeps resource requirements manageable. This balance is key for deployment on standard hardware infrastructures.
Experts in the specified domains contributed to data curation to ensure relevance and accuracy. This collaborative effort improves the model's understanding of specialized terminology and practices. The result is better performance on domain specific queries.
How does the Hunyuan Hy4 preview perform according to internal benchmarks?
An internal blind evaluation involved 163 experts assessing the model on 203 engineering tasks. The average score achieved was 2.99 out of 4.00. This places the model in a competitive position relative to other frontier models.
GLM 5.3 received a score of 2.92 out of 4.00 in the same evaluation. The win rate for Hy4 preview against it included 46.8 percent wins 12.8 percent ties and 40.4 percent losses. These figures indicate a slight edge in the tested scenarios.
Kimi K3 scored 2.94 out of 4.00. Against this model the Hy4 preview had 51.2 percent wins 7.9 percent ties and 40.9 percent losses. The results highlight consistent performance across different comparison points.
The number of experts involved in the evaluation lends credibility to the results. 203 tasks provide a broad sample of engineering challenges. Blind evaluation prevents bias in scoring.
Comparison with GLM 5.3 and Kimi K3 shows the model is in the same league. Slight advantages in average score can translate to noticeable differences in user experience. The win loss breakdown provides detailed insight into relative strengths.
Scores in this range indicate strong capability but room for improvement in some tasks. The win rates show that the model often outperforms but not always. Ties suggest similar performance levels in certain cases.
| Model | Total Parameters | Active Parameters | Context Window | Average Score |
|---|---|---|---|---|
| Hunyuan Hy4 preview | 770B | 49B | Exceeding 1M tokens | 2.99/4.00 |
| GLM-5.3 | Not specified | Not specified | Not specified | 2.92/4.00 |
| Kimi K3 | Not specified | Not specified | Not specified | 2.94/4.00 |
The benchmarks focus on areas where performance gains are most valuable. These include long horizon software engineering tasks. Document heavy office work also benefits from the model's capabilities. Scientific research applications see improvements due to the large context window.
What are the pricing details and access methods for users?
The API pricing per 1 million tokens is 0.834 dollars for input. Output tokens cost 2.501 dollars per million. Cached input is available at 0.042 dollars per million. This pricing is designed to be competitive for frequent usage.
The cached input price of 0.042 dollars encourages use of context caching for repeated queries. This feature reduces costs for applications with recurring patterns. It is a practical consideration for developers optimizing expenses.
Per million token pricing allows easy estimation of costs for projects of varying sizes. Enterprises can budget accordingly based on expected usage volumes. The structure supports both small scale testing and large scale production.
Users can access the model through Tencent Cloud TokenHub. OpenRouter provides an alternative entry point for developers. The model is integrated into WorkBuddy for general productivity. CodeBuddy offers coding assistance features powered by the model.
Free access to the Hy4 preview on WorkBuddy and CodeBuddy is offered for two weeks at launch. The free Hy3 access has been extended until September 30 2026. These promotions encourage users to test the new capabilities without initial cost.
- API through Tencent Cloud TokenHub
- API through OpenRouter
- Integration in WorkBuddy for productivity tasks
- Integration in CodeBuddy for development work
- Integration in Yuanbao and ima platforms
What implications does this release have for the market and stakeholders?
The open source nature allows the community to build upon the model. This contributes to the open source frontier in AI development. Cost effective pricing supports adoption in cost sensitive environments.
Open sourcing the model follows a trend in the industry toward greater accessibility. This can accelerate innovation as researchers experiment with the weights. The 770B scale provides a substantial base for fine tuning by the community.
Affordable pricing enables smaller organizations to utilize frontier level models. This democratizes access to high performance AI for real world tasks. Competition on price may lead to better options for consumers overall.
The presence in multiple apps means the model can be used across different professional contexts. From coding to document management the utility is broad. This multi platform approach maximizes the impact of the release.
Stakeholders in the software industry gain from the engineering focused training data. Finance and security sectors benefit from specialized knowledge embedded in the model. Gaming applications can leverage the capabilities for complex simulations and interactions.
The integration into existing Tencent tools facilitates seamless adoption by current users. This reduces the barrier for enterprises looking to incorporate advanced AI into workflows. The model competes directly with GLM-5.3 and Kimi K3 on performance and cost metrics.
The release may prompt other companies to accelerate their own open source efforts. This could lead to a more vibrant ecosystem of models. Users benefit from increased choice and innovation.
For AI agents and retrieval augmented generation systems the long context is particularly advantageous. It reduces the need for complex chunking strategies in RAG setups. This can improve accuracy in knowledge intensive applications.
What have experts stated about the capabilities of the Hunyuan Hy4 preview?
The development team has provided insights into the model's strengths. The focus remains on areas that present the greatest challenges for AI systems.
Today we're releasing Hy4 preview our most capable model to date. It's a 770B parameter model with 49B active parameters and a 1M token context window and it makes the biggest gains where the work is hardest long horizon software engineering document heavy office work and scientific research.Tencent Hy team Research and development
What can be expected in the coming period after the launch?
The two week free period will allow extensive testing by users. Extended support for the previous model ensures continuity. Further integrations may expand the reach of the technology.
Continued development may focus on refining the active parameter efficiency. The current 49B active parameters already provide a balance between performance and speed.
The launch date of August 28 2026 marks an important milestone for Tencent in the AI space. It demonstrates ongoing investment in frontier model development. Observers will watch for subsequent releases in the series.
Frequently asked
What is the context window size of the Hunyuan Hy4 preview?
The context window of the Hunyuan Hy4 preview exceeds 1 million tokens. This size allows processing of extensive inputs in tasks such as document analysis and long horizon planning. The feature supports complex workflows without loss of coherence.
Sources
- Tencent — Tencent has released and open-sourced Tencent Hy4 preview, a next-generation large language model with 770B total parameters and 49B active parameters, and a context window exceeding 1M tokens.
- Tencent — Today we're releasing Hy4 preview, our most capable model to date. It's a 770B-parameter model with 49B active parameters and a 1M-token context window, and it makes the biggest gains where the work is hardest: long-horizon software engineering, document-heavy office work, and scientific research.
- Hugging Face — Hy4 preview is a new generation mixed expert (MoE) flagship model developed by the Tencent Hunyuan team. The total parameter amount of the model is 770B, and 49B is activated per token.