# b.ai Offers Google Gemini 3.6 Flash and 3.5 Flash-Lite via API for Agent Development

> The platform provides immediate access to the new models with large context windows and tool support, enabling developers to scale AI agents efficiently without additional setup.

*Published 2026-07-29 · By Marcus Vance*

b.ai is a fast-follower API provider delivering Google's Gemini 3.6 Flash and Gemini 3.5 Flash-Lite to developers matching official specifications and pricing for agentic and multimodal workloads.

The integration of these models on the b.ai platform represents a strategic move to cater to the needs of developers working on AI agents. With the ability to handle large amounts of data through the extensive context window, these models can process complex queries that involve multiple types of media. This is particularly useful for applications in areas such as automated content creation, data analysis, and interactive systems. The immediate availability ensures that users can start building and testing their applications right away without any lag in access to the latest technology from Google.

## What background information is available on the development of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite?

The Gemini Flash series has been created to meet the demands of production environments where AI agents are deployed at scale. According to statements from Google, the focus is on providing higher token efficiency, lower latency, and more reliable performance. This is essential for workflows that require consistent operation over long periods. The models build upon the foundation of earlier Gemini versions to offer enhanced capabilities in multimodal processing and tool use. The emphasis on the sweet spot of efficiency and quality allows for better scaling of agentic workflows across different industries.

Stakeholders in the AI field have been calling for models that can handle the rigors of real-world use cases. The Flash models address this by incorporating features that reduce resource consumption while maintaining high levels of performance. This development is part of a larger trend in the industry to optimize AI for practical applications rather than just increasing model size. The release also highlights the competitive landscape where providers like b.ai play a crucial role in distribution.

## What new features and capabilities are present in the latest Gemini models?

The new models come equipped with a range of tools that facilitate advanced agent behaviors. Function calling allows the models to interact with external systems in a structured manner. Structured outputs ensure that the responses are in a format that can be easily parsed by other software components. Code execution capability enables the models to run code snippets directly, which is invaluable for coding agents. Grounding and URL context help in providing accurate information based on reliable sources. File-search workflows support the handling of large document sets.

Multimodal support extends to various input types, allowing the models to analyze images, videos, audio clips, and PDF documents alongside text. This broad input capability opens up possibilities for applications that require understanding of diverse data formats. The output remains text-based, which is standard for most agent interactions. The large context window permits the inclusion of extensive conversation histories or large datasets in a single prompt.

## How do the technical specifications of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite compare?

Technical specifications comparison between the two new Gemini models available on b.aiModelContext WindowMax Output TokensOutput Tokens per SecondPrimary Use CaseGemini 3.6 Flash1,048,576 tokens65,536 tokensNot specifiedAgentic coding and multimodal knowledge workGemini 3.5 Flash-Lite1,048,576 tokens65,536 tokens350Low-latency, high-throughput multimodal processing

## What performance improvements have been observed in these models?

Data from the Artificial Analysis Index indicates that Gemini 3.6 Flash uses 17% less output tokens than the previous version in general cases. In specific benchmarks such as DeepSWE, the reduction can be as high as 65%. These improvements contribute to lower costs and faster execution times for users running multiple queries. The efficiency gains are particularly beneficial for long-running agent sessions where token usage can accumulate quickly.

The speed of Gemini 3.5 Flash-Lite at 350 output tokens per second allows for quick generation of responses in time-sensitive applications. This performance metric is crucial for user-facing AI agents that need to provide instant feedback. The combination of large context and high speed makes these models versatile for a variety of use cases.

## What implications does this have for the market and various stakeholders?

The availability through b.ai means that developers have an additional option for accessing these models, potentially at competitive rates since pricing matches the official. This could lead to increased innovation as more teams experiment with agentic systems. Stakeholders such as enterprise customers may find it easier to integrate these models into their existing infrastructure due to the comprehensive tool support.

The market may see a shift towards more specialized API providers that act as intermediaries, offering convenience and immediate access. This model of distribution can help in democratizing access to frontier models. For Google, it extends the reach of their models through partners like b.ai.

## What expert reactions have been recorded regarding the Gemini model releases?

Tulsee Doshi, speaking on behalf of the Gemini team, has highlighted the models' suitability for building AI agents at scale. The statement focuses on the efficiency, latency, and reliability aspects that are critical for production use. This reflects the priorities in the current AI development cycle.

> Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.Tulsee Doshi, Senior Director, Product Management, on behalf of the Gemini team

Other experts in the field have echoed the importance of such advancements for the growth of the AI agent ecosystem. The focus on agentic workloads indicates where the industry is heading in terms of practical AI applications.

## What can be expected next in the development and deployment of these models?

Future updates may include refinements based on user feedback and new benchmark results. b.ai is likely to continue its role in providing quick access to updates from Google. The ongoing evolution of these models will likely include improvements in other areas such as accuracy and additional tool integrations.

- Continued optimization of token efficiency based on real-world usage data.
- Expansion of multimodal capabilities to include more input formats.
- Increased focus on security and grounding features for enterprise adoption.
- Release of updated performance metrics from sources like Artificial Analysis.
- Potential introduction of hybrid models combining features from both variants.

The collaboration between b.ai and Google in this manner sets a precedent for how frontier models can be distributed more widely. This approach benefits developers by reducing the time to market for new AI capabilities. As the technology matures, more such partnerships are expected to emerge in the AI space.

## Sources

1. [Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows.](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)
2. [Gemini 3.6 Flash is a Gemini 3-series model for agentic coding, multimodal knowledge work, spatial reasoning, and multi-step workflows. On B.AI, use the model ID gemini-3.6-flash. ... Context Window: Supports up to 1,048,576 input tokens. Max Output: Supports up to 65,536 output tokens. ... Multimodal Input: Supports text, image, video, audio, and PDF input with text output. Tool Integration: Supports function calling, structured outputs, code execution...](https://docs.b.ai/llmservice/models/gemini-3-6-flash/)
3. [Gemini 3.5 Flash-Lite is a Gemini 3-series model for low-latency, high-throughput multimodal processing. ... On B.AI, use the model ID gemini-3.5-flash-lite. ... Context Window: Supports up to 1,048,576 input tokens. Max Output: Supports up to 65,536 output tokens. ... Multimodal Input: Supports text, image, video, audio, and PDF input with text output. Tool Integration: Supports function calling, structured outputs, code execution...](https://docs.b.ai/llmservice/models/gemini-3-5-flash-lite/)

---
Source: https://aiintelreport.com/frontier-models/b-ai-offers-google-gemini-3-6-flash-and-3-5-flash-lite-via-api
Index: https://aiintelreport.com/llms.txt · Full text: https://aiintelreport.com/llms-full.txt
