# Kimi K3 by Moonshot AI Launches on Amazon Bedrock with 1M Context and Prompt Caching

> The 2.8 trillion parameter open-weight model becomes the first of its kind to support explicit prompt caching on the platform, offering developers new options for long context applications at $3 input and $15 output per million tokens.

*Published 2026-09-19 · By Marcus Vance*

Kimi K3 is Moonshot AI's most capable open-weight model, the first open model to reach 2.8 trillion parameters, with a 1-million-token context window, native vision capabilities, and explicit prompt caching support.

The launch of Kimi K3 on Amazon Bedrock provides a new avenue for accessing one of the largest open-weight models available today. Moonshot AI has made its flagship model available through AWS managed infrastructure. This allows companies to utilize the capabilities without the need for extensive hardware resources. The general availability date of September 18, 2026, marks the point where developers can begin integrating the model into their workflows. The 2.8 trillion parameter count sets a new benchmark for open models. Users benefit from the 1 million token context window which enables processing of very long documents or codebases. Native vision support adds the ability to analyze images alongside text. The introduction of explicit prompt caching is a notable first for open-weight models on this platform. This feature can help reduce costs and latency for repeated queries by storing checkpoints of at least 1,024 tokens with a minimum thirty minute time to live. The pricing is listed as three dollars for input and fifteen dollars for output per million tokens. Cache operations have separate rates with reads at thirty cents and writes at three dollars seventy five cents for the thirty minute duration. These elements combine to make the model suitable for frontier level tasks in coding and reasoning.

Background on the development of such large models shows the rapid progress in the field. Moonshot AI has focused on scaling parameters while maintaining open weights. The Kimi K3 model builds on previous iterations with advanced attention mechanisms. The Kimi Delta Attention and Attention Residuals are key to handling the scale. This allows the model to achieve high performance on complex tasks. The context window of one million tokens is among the largest offered in open models. It opens possibilities for applications that require understanding entire books or large code repositories in one go. Vision capabilities mean the model can process visual data natively without additional components. The availability on Bedrock integrates it into an ecosystem used by many enterprises. This could lead to wider adoption compared to self-hosted versions. The explicit prompt caching is designed to optimize for scenarios where prompts are reused over time. Organizations can save on compute costs by caching common prefixes or instructions. The minimum requirements ensure that only substantial checkpoints are stored to maintain efficiency.

## What technical features set Kimi K3 apart from other models on Amazon Bedrock?

The technical architecture of Kimi K3 includes specific innovations that support its parameter count and context length. Built on Kimi Delta Attention and Attention Residuals, the model can handle the computational demands of 2.8 trillion parameters. This design choice contributes to its ability to perform frontier intelligence tasks. The 1 million token context window requires sophisticated memory management during inference. Native vision capabilities allow direct input of images for analysis in conjunction with text prompts. The model supports multiple modalities including image and text as standard. Explicit prompt caching is implemented with specific parameters to ensure usability. Checkpoints must contain at least 1,024 tokens to be eligible for caching. The time to live is set at a minimum of thirty minutes to balance storage and relevance. These features are documented in the model card provided by AWS. The inference IDs allow selection of global or US specific profiles for optimized routing. This technical setup positions Kimi K3 as a versatile tool for developers working on long horizon projects.

## What are the market and stakeholder implications of Kimi K3 on Bedrock?

The release has implications for the competitive landscape in AI services. By bringing an open large model to Bedrock, AWS expands its offerings in the frontier models category. Stakeholders in enterprise AI can now access this scale through familiar interfaces. This may influence decisions on whether to use open or closed models. The pricing structure makes it accessible for high volume use cases. Developers in coding and knowledge work fields stand to benefit from the context and vision features. The prompt caching can lead to more efficient application designs. Market analysts may see this as a move to democratize access to large models. It could encourage other providers to offer similar features. The open weight nature allows for customization if needed though on Bedrock it is managed. Overall, it strengthens the position of managed platforms in the AI ecosystem.

## What reactions have come from the launch of the model?

> Today, we are introducing Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.Moonshot AI

The company statement from Moonshot AI underscores the model's intended use cases in long horizon coding and reasoning. This aligns with the technical specifications made available through the Bedrock integration. Users can now test these capabilities in a managed environment. The emphasis on being the first open 3T-class model highlights the parameter scale achieved. Reactions from the ecosystem are likely to center on the combination of scale and new features like prompt caching. The availability through established inference profiles supports immediate experimentation. Organizations evaluating frontier models will note the pricing and caching options as differentiators. The launch provides a concrete example of how open models can be deployed at scale via cloud services.

## What developments are anticipated following the Kimi K3 release on Bedrock?

Following the general availability, further updates to the model or additional features may be introduced by Moonshot AI and AWS. The current capabilities provide a strong foundation for immediate use in production environments. Stakeholders will monitor performance metrics and user feedback to guide future iterations. The open weight aspect may lead to community contributions that enhance the model over time. Integration with other AWS services could expand the use cases. The focus on long context and vision suggests continued emphasis on multimodal applications. Pricing may be adjusted based on demand and operational costs. The success of prompt caching could inspire similar implementations for other models on the platform. Overall, the launch sets the stage for more advanced offerings in the frontier models space.

The 1 million token context window allows the model to maintain coherence over extremely long inputs. This is particularly useful in scenarios involving extensive code bases where the model can reference multiple files at once. In knowledge work, it can process full reports or datasets in a single prompt. The native vision means images can be described or analyzed within the same context. This multimodal approach reduces the need for separate vision models. The explicit prompt caching complements this by allowing reuse of common context elements. For instance, a system prompt or reference material can be cached and reused across sessions. This reduces the effective cost for applications with repetitive structures. The minimum token requirement ensures that only meaningful amounts are cached. The thirty minute TTL means caches are temporary and relevant to current tasks. These mechanisms together enhance the practicality of using such a large model in real world applications.

In the area of long-horizon coding, Kimi K3 can handle complex projects that span thousands of lines of code. The large context allows the model to understand dependencies across the entire project. This can improve the quality of code generation and debugging suggestions. Developers can provide the full repository as context without truncation. The vision capabilities can be used to analyze diagrams or screenshots of code issues. The prompt caching can store common coding patterns or library references. This leads to faster iteration cycles in software development. Enterprises may adopt the model for internal tools that require deep understanding of their codebases. The open weight nature allows inspection of the model if needed for compliance. The Bedrock integration provides the security and scalability expected from AWS services. This combination makes it attractive for regulated industries. The pricing is structured to support both experimentation and production use.

Compared to other models available on Amazon Bedrock, Kimi K3 stands out due to its parameter scale and context length. It is the first open model to offer the 2.8 trillion parameter size. The prompt caching feature is unique among open-weight options on the platform. The pricing is specified separately for input and output to reflect the computational demands. The support for multiple modalities expands its utility beyond text only models. The inference profiles allow for global or regional deployment to meet latency requirements. This flexibility is important for international users. The general availability means no waitlist or preview restrictions apply. Users can immediately begin testing the model through the Bedrock console or APIs. The documentation provides detailed guidance on configuration and best practices. This transparency aids in adoption by technical teams. The launch reflects the trend toward larger and more capable open models in the industry.

The release also highlights the role of cloud providers in distributing advanced AI models. AWS has facilitated the deployment of Kimi K3 through its infrastructure. This reduces the operational burden on users who would otherwise need to manage large scale hardware. The model card details the supported features and limitations. It includes information on the pricing and the requirements for prompt caching. Organizations can use this information to plan their AI strategies. The availability may influence decisions on model selection for specific use cases. For tasks requiring long context, Kimi K3 offers a competitive option. The explicit caching can be a deciding factor for cost sensitive applications. Overall, the integration contributes to the diversity of models available on the platform. This benefits users by providing choices tailored to their needs.

Expert reactions to the launch have focused on the scale and features offered. The company statement emphasizes the model's design for frontier intelligence. It highlights the use of specific attention mechanisms to achieve the parameter count. The statement also notes the applications in coding, knowledge work, and reasoning. This aligns with the capabilities provided through Bedrock. Users are expected to explore these areas in their implementations. The combination of open weights and managed service creates new opportunities. Future updates may build on this foundation with improved performance or additional features. The current offering sets a high standard for open models on cloud platforms. Monitoring the adoption rates will provide insights into its impact on the market. The launch is a step forward in making frontier models more accessible.

The overall impact of this model release extends to the broader AI research community. By making a 2.8 trillion parameter model openly available through a major platform, it encourages further innovation. Researchers can study the model to understand scaling laws at this size. The features like prompt caching provide practical tools for efficient inference. This can lead to new research on optimization techniques. The vision support opens avenues for multimodal research. The context window size invites studies on long term memory in language models. The pricing model allows for experimentation at scale. As more users adopt the model, feedback will likely drive improvements. The partnership between Moonshot AI and AWS demonstrates how collaborations can accelerate deployment. This sets a precedent for future model releases on the platform. The availability on September 18, 2026, will be remembered as a key date in the timeline of open frontier models.

- Select the model using one of the inference IDs such as global.moonshotai.kimi-k3.
- Ensure prompts for caching meet the minimum of 1,024 tokens per checkpoint.
- Set the time to live parameter to at least thirty minutes for the cache.
- Submit the initial prompt to create the cache entry.
- Utilize cache read operations for follow up queries to reduce costs.

Pricing structure for Kimi K3 inference on Amazon Bedrock from AWS documentationInference ProfileInput Price (per million tokens)Output Price (per million tokens)Cache Read PriceCache Write Price (30 min)global.moonshotai.kimi-k3$3.00$15.00$0.30$3.75us.moonshotai.kimi-k3$3.00$15.00$0.30$3.75

## Sources

1. [Today, Kimi K3 from Moonshot AI is generally available on Amazon Bedrock... Kimi K3 is the first open weight model on Amazon Bedrock to support explicit prompt caching](https://aws.amazon.com/about-aws/whats-new/2026/09/moonshot-ai-kimi-k3-on-amazon-bedrock/)
2. [Kimi K3 is Moonshot AI's most capable open-weight model... Pricing table for Global CRIS and US CRIS; model IDs global.moonshotai.kimi-k3 and us.moonshotai.kimi-k3; supports explicit prompt caching](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html)
3. [Today, Kimi K3 from Moonshot AI is available on Amazon Bedrock... According to Moonshot AI, Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters... first open-weight model on Amazon Bedrock to support explicit prompt caching](https://aws.amazon.com/blogs/machine-learning/introducing-kimi-k3-on-amazon-bedrock/)
4. [Company statement on Kimi K3 capabilities](https://www.kimi.ai/blog/kimi-k3)

---
Source: https://aiintelreport.com/frontier-models/kimi-k3-moonshot-ai-amazon-bedrock-availability
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
