Enterprise AI
Kimi K3 Launches as Moonshot AI's Open 3T-Class MoE with 1M Context
Moonshot AI delivers live API access to its 2.8T-parameter model and schedules full weights release for July 27, 2026, offering enterprises new options for large-context and multimodal workloads.
Kimi K3 is a 2.8T-parameter MoE model with native vision capabilities and a 1-million-token context window.
The introduction of Kimi K3 marks a significant development in the landscape of large language models available to enterprises. With its scale of 2.8 trillion parameters, the model offers substantial capacity for complex reasoning tasks that are common in business environments such as financial analysis and customer service automation. Enterprises have long sought models that can handle extensive contexts without losing coherence, and the 1 million token window addresses this need directly by allowing entire datasets or long conversations to be processed in one go. This capability can lead to more accurate outputs in applications like contract review, where missing context from chunking can introduce errors. The native vision capabilities further extend its utility to multimodal tasks, enabling analysis of documents containing images and charts without separate preprocessing pipelines. As a result, organizations can streamline their AI workflows and potentially achieve better return on investment through reduced integration complexity.
Enterprises are increasingly turning to large context models to handle complex tasks that require understanding of lengthy inputs. Kimi K3's design directly addresses this by providing the capacity to process up to one million tokens in a single inference pass. This reduces the need for retrieval augmented generation techniques in some cases, simplifying architecture and potentially improving response quality. The inclusion of native vision opens doors for applications in industries like healthcare, where both text reports and medical images need to be analyzed together. With the model already live on the API, companies can begin integrating it into their systems immediately while preparing for the open weights release that will allow for greater customization and control.
What background informs the development of Kimi K3?
Moonshot AI has positioned itself as a key player in advancing open models that compete with closed systems. The company has built Kimi K3 on proprietary techniques including Kimi Delta Attention and Attention Residuals, which contribute to its efficiency despite the large parameter count. These innovations allow the model to maintain performance while managing the computational demands of a 3T-class architecture. In the context of enterprise AI, such developments are critical because they provide alternatives to models from US-based companies like those behind Claude Fable 5 and GPT 5.6 Sol, which remain closed and may raise concerns around data sovereignty and access restrictions. The open nature planned for Kimi K3 could allow enterprises to deploy the model on their own infrastructure, mitigating risks associated with sending sensitive data to third-party APIs.
The escalation in open weights releases reflects broader industry shifts where organizations prioritize control over their AI assets. By releasing a model of this scale, Moonshot AI contributes to a more diverse set of choices for chief AI officers tasked with balancing performance, cost, and compliance. Historical patterns show that open models often accelerate adoption in regulated sectors because they permit auditing and modification. Kimi K3 follows this trajectory with its combination of scale and planned openness, creating pressure on closed frontier providers to respond with improved transparency or pricing.
What new features does Kimi K3 introduce in detail?
Kimi K3 stands out as the first open model in the 3T-class category to incorporate both a 1 million token context and native vision. The architecture relies on Stable LatentMoE, which activates only 16 out of 896 experts during inference. This sparse activation helps in managing the inference costs and latency, making it more practical for enterprise scale deployments where high throughput is required. The model is already integrated into Kimi.com for general use, Kimi Work for productivity tools, Kimi Code for development assistance, and the dedicated Kimi API for programmatic access. This multi-platform availability ensures that different teams within an organization can leverage the model according to their specific needs, from coding assistance to document processing.
What are the technical specifics of the model architecture?
The technical foundation of Kimi K3 includes several advanced components that enable its performance. Kimi Delta Attention is designed to enhance the attention mechanism for better handling of long sequences, which is essential for the 1 million token context. Attention Residuals provide additional stability during training and inference, reducing the likelihood of performance degradation in deep networks. The Stable LatentMoE component allows for efficient expert selection, activating a small subset of the total experts to process each token. With 896 experts in total but only 16 active, the model achieves a balance between capacity and efficiency. Native vision means the model can directly process image inputs alongside text, which is useful for enterprises dealing with visual data such as scanned documents or product images in e-commerce applications.
| Feature | Specification |
|---|---|
| Parameters | 2.8 trillion |
| Architecture | MoE with Stable LatentMoE |
| Context Window | 1 million tokens |
| Experts Total | 896 |
| Experts Activated | 16 |
| Capabilities | Native vision, text |
| Release Status | API live, weights July 27, 2026 |
How does the API pricing structure work for Kimi K3?
The pricing for the Kimi K3 API has been set to balance accessibility with the high computational requirements of the model. Input tokens are charged at $3.00 per million for cache-miss cases, while cache-hit inputs are significantly lower at $0.30 per million. Output tokens cost $15.00 per million. This tiered approach incentivizes the use of caching mechanisms to reduce expenses for repeated queries, which is common in enterprise settings where similar prompts may be used across multiple users or sessions. For organizations with high volume usage, the cache-hit rate can lead to substantial savings, making the model more viable for production workloads.
What are the market and stakeholder implications?
The release of Kimi K3 has implications for various stakeholders in the AI ecosystem. For enterprises, it offers a new option for building custom AI solutions with the potential for self-hosting once weights are released. This can enhance data sovereignty by keeping proprietary information within organizational boundaries rather than relying on external providers. Developers gain access to a high-capacity model for experimentation and integration into applications. The competition with closed models from US frontiers may drive down prices across the industry and accelerate innovation in open source alternatives. However, the high output pricing may require careful cost management in applications that generate long responses.
- Evaluate caching strategies to minimize input costs.
- Assess use cases for the 1 million token context in document-heavy workflows.
- Plan for self-hosting options after the July 27 release to address data sovereignty.
- Compare performance against existing models in Claude Fable 5 and GPT 5.6 Sol categories.
- Monitor capacity limits as demand has already stressed infrastructure.
What expert reactions have emerged following the launch?
Reactions from the community and officials at Moonshot AI highlight the strong interest in Kimi K3. The demand has been notable, indicating a market readiness for such advanced open models.
Kimi K3 has received far more love than we expected, and our GPUs are feeling it. Over the past 48 hours, demand has pushed close to the limits of our current capacity.Kimi_Moonshot (Moonshot AI)
What comes next in the open weights escalation?
With the full weights scheduled for release on July 27, 2026, the landscape for enterprise AI is poised for further change. Organizations will have the opportunity to fine-tune and deploy Kimi K3 on private infrastructure, potentially integrating it with on-premises systems for enhanced security. This move could intensify the shift toward open models, pressuring closed providers to offer more transparent or accessible options. Enterprises should prepare by testing the API now to understand performance characteristics before committing to self-hosted versions. The combination of large context, vision, and eventual open access positions Kimi K3 as a tool that can support complex, multimodal enterprise applications while offering flexibility in deployment.
Furthermore, the availability through multiple interfaces means that non-technical stakeholders can also benefit from the model through user-friendly platforms like Kimi Work. This broadens the adoption potential beyond AI specialists to include business analysts and operations teams. In terms of implementation, the model design choices around expert activation contribute to more predictable resource usage, which is a key consideration for budgeting AI infrastructure. As more details emerge post-release, enterprises will need to conduct thorough evaluations to determine the best fit for their specific requirements and compliance standards.
How does this affect AI strategy for chief AI officers?
Chief AI officers evaluating Kimi K3 must consider the balance between capability and cost. The high output price suggests that applications should be optimized for concise responses where possible. At the same time, the context length can reduce the number of API calls needed for long inputs, offsetting some costs. The upcoming open release allows for experimentation with quantization or distillation to further optimize for enterprise hardware. Organizations that invest early in integration testing stand to gain advantages in both performance and strategic positioning against competitors still reliant on closed models.
Overall, this development underscores the rapid pace of progress in open AI models and the importance of monitoring releases from companies like Moonshot AI to stay competitive in the enterprise space. Decision makers should track updates on capacity and performance benchmarks as the model scales in usage.
Frequently asked
When will the full weights for Kimi K3 be available?
The full model weights will be released by July 27, 2026.
What is the context window size for Kimi K3?
Kimi K3 supports a 1-million-token context window.
How is Kimi K3 accessed today?
Kimi K3 is available today via Kimi.com, Kimi Work, Kimi Code, and the Kimi API.
Sources
- Moonshot AI — Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. ... The full model weights will be released by July 27, 2026. ... Pricing is $0.30/MTok for cache-hit input, $3.00/MTok for cache-miss input, and $15.00/MTok for output.
- Moonshot AI — Kimi K3 has launched! ... Cache Hit $0.30 / MTok Input $3.00 / MTok Output $15.00 / MTok
- Moonshot AI — Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. ... The full model weights will be released by July 27, 2026.
- X — Demand quote for Kimi K3