Tuesday, July 28, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

b.ai Offers Google Gemini 3.6 Flash and 3.5 Flash-Lite via API for Agent Development

The platform provides immediate access to the new models with large context windows and tool support, enabling developers to scale AI agents efficiently without additional setup.

6 MIN READ
A realistic live-action photograph of an anonymous developer seated with back to the viewer in a modern professional technology office workspace featuring a large wooden desk covered in computer hardware including open laptops with screens displaying abstract node diagrams and flow charts for AI agent architectures external solid state drives connected via thick black cables bundles of multicolored Ethernet cords linking to a tall metal server rack positioned against the far wall a tablet device resting on the desk surface showing grid based data visualizations and performance metrics charts ergonomic office chair with mesh backrest positioned in front of the desk keyboard mouse and headset peripherals neatly arranged on the desk surface potted green plants on the windowsill behind the desk floor to ceiling windows revealing a distant city skyline outside tall bookshelves along the side walls holding rows of technical reference volumes a second anonymous figure standing near the server rack adjusting connections on the hardware units a whiteboard mounted on the wall with hand drawn abstract diagrams of system architectures but no readable markings coffee mug and stainless steel water bottle placed beside the laptops on the desk surface additional peripherals including webcams microphones and external monitors all showing non textual graphical interfaces this entire setup represents the direct API based access to advanced large context AI models from major providers enabling efficient scaling of agent development tools and integrations without any further configuration steps or installations required the workspace environment includes visible network infrastructure cables running along the floor to connect multiple devices soft natural daylight illuminating the room from the windows neutral colored walls and carpeted flooring with a focus on the hardware and software development tools that support immediate platform access for building and deploying intelligent agents at scale the scene emphasizes real world elements such as the server hardware representing cloud infrastructure the laptops and monitors illustrating developer interfaces for model interactions and the overall office setting conveying a productive environment for technology professionals working with API offerings from established AI model providers the composition centers on the desk and server area to highlight the seamless integration of tool support and context handling capabilities in practical developer workflows.
Illustration: AI Intel Report

b.ai is a fast-follower API provider delivering Google's Gemini 3.6 Flash and Gemini 3.5 Flash-Lite to developers matching official specifications and pricing for agentic and multimodal workloads.

The integration of these models on the b.ai platform represents a strategic move to cater to the needs of developers working on AI agents. With the ability to handle large amounts of data through the extensive context window, these models can process complex queries that involve multiple types of media. This is particularly useful for applications in areas such as automated content creation, data analysis, and interactive systems. The immediate availability ensures that users can start building and testing their applications right away without any lag in access to the latest technology from Google.

What background information is available on the development of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite?

The Gemini Flash series has been created to meet the demands of production environments where AI agents are deployed at scale. According to statements from Google, the focus is on providing higher token efficiency, lower latency, and more reliable performance. This is essential for workflows that require consistent operation over long periods. The models build upon the foundation of earlier Gemini versions to offer enhanced capabilities in multimodal processing and tool use. The emphasis on the sweet spot of efficiency and quality allows for better scaling of agentic workflows across different industries.

Stakeholders in the AI field have been calling for models that can handle the rigors of real-world use cases. The Flash models address this by incorporating features that reduce resource consumption while maintaining high levels of performance. This development is part of a larger trend in the industry to optimize AI for practical applications rather than just increasing model size. The release also highlights the competitive landscape where providers like b.ai play a crucial role in distribution.

What new features and capabilities are present in the latest Gemini models?

The new models come equipped with a range of tools that facilitate advanced agent behaviors. Function calling allows the models to interact with external systems in a structured manner. Structured outputs ensure that the responses are in a format that can be easily parsed by other software components. Code execution capability enables the models to run code snippets directly, which is invaluable for coding agents. Grounding and URL context help in providing accurate information based on reliable sources. File-search workflows support the handling of large document sets.

Multimodal support extends to various input types, allowing the models to analyze images, videos, audio clips, and PDF documents alongside text. This broad input capability opens up possibilities for applications that require understanding of diverse data formats. The output remains text-based, which is standard for most agent interactions. The large context window permits the inclusion of extensive conversation histories or large datasets in a single prompt.

How do the technical specifications of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite compare?

Technical specifications comparison between the two new Gemini models available on b.ai
ModelContext WindowMax Output TokensOutput Tokens per SecondPrimary Use Case
Gemini 3.6 Flash1,048,576 tokens65,536 tokensNot specifiedAgentic coding and multimodal knowledge work
Gemini 3.5 Flash-Lite1,048,576 tokens65,536 tokens350Low-latency, high-throughput multimodal processing

What performance improvements have been observed in these models?

Data from the Artificial Analysis Index indicates that Gemini 3.6 Flash uses 17% less output tokens than the previous version in general cases. In specific benchmarks such as DeepSWE, the reduction can be as high as 65%. These improvements contribute to lower costs and faster execution times for users running multiple queries. The efficiency gains are particularly beneficial for long-running agent sessions where token usage can accumulate quickly.

The speed of Gemini 3.5 Flash-Lite at 350 output tokens per second allows for quick generation of responses in time-sensitive applications. This performance metric is crucial for user-facing AI agents that need to provide instant feedback. The combination of large context and high speed makes these models versatile for a variety of use cases.

What implications does this have for the market and various stakeholders?

The availability through b.ai means that developers have an additional option for accessing these models, potentially at competitive rates since pricing matches the official. This could lead to increased innovation as more teams experiment with agentic systems. Stakeholders such as enterprise customers may find it easier to integrate these models into their existing infrastructure due to the comprehensive tool support.

The market may see a shift towards more specialized API providers that act as intermediaries, offering convenience and immediate access. This model of distribution can help in democratizing access to frontier models. For Google, it extends the reach of their models through partners like b.ai.

What expert reactions have been recorded regarding the Gemini model releases?

Tulsee Doshi, speaking on behalf of the Gemini team, has highlighted the models' suitability for building AI agents at scale. The statement focuses on the efficiency, latency, and reliability aspects that are critical for production use. This reflects the priorities in the current AI development cycle.

Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.Tulsee Doshi, Senior Director, Product Management, on behalf of the Gemini team

Other experts in the field have echoed the importance of such advancements for the growth of the AI agent ecosystem. The focus on agentic workloads indicates where the industry is heading in terms of practical AI applications.

What can be expected next in the development and deployment of these models?

Future updates may include refinements based on user feedback and new benchmark results. b.ai is likely to continue its role in providing quick access to updates from Google. The ongoing evolution of these models will likely include improvements in other areas such as accuracy and additional tool integrations.

  1. Continued optimization of token efficiency based on real-world usage data.
  2. Expansion of multimodal capabilities to include more input formats.
  3. Increased focus on security and grounding features for enterprise adoption.
  4. Release of updated performance metrics from sources like Artificial Analysis.
  5. Potential introduction of hybrid models combining features from both variants.

The collaboration between b.ai and Google in this manner sets a precedent for how frontier models can be distributed more widely. This approach benefits developers by reducing the time to market for new AI capabilities. As the technology matures, more such partnerships are expected to emerge in the AI space.

Frequently asked

How can developers access Gemini 3.6 Flash on b.ai?

Developers access Gemini 3.6 Flash on b.ai using the model ID gemini-3.6-flash. The model supports the full set of specifications including the 1,048,576-token context window and multimodal inputs as documented in the b.ai model documentation.

What is the primary difference between Gemini 3.6 Flash and Gemini 3.5 Flash-Lite?

Gemini 3.6 Flash focuses on efficiency gains such as reduced token usage for agentic and coding tasks while Gemini 3.5 Flash-Lite prioritizes low latency and high throughput for multimodal processing at 350 output tokens per second.

Sources

  1. Google — Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows.
  2. B.AI — Gemini 3.6 Flash is a Gemini 3-series model for agentic coding, multimodal knowledge work, spatial reasoning, and multi-step workflows. On B.AI, use the model ID gemini-3.6-flash. ... Context Window: Supports up to 1,048,576 input tokens. Max Output: Supports up to 65,536 output tokens. ... Multimodal Input: Supports text, image, video, audio, and PDF input with text output. Tool Integration: Supports function calling, structured outputs, code execution...
  3. B.AI — Gemini 3.5 Flash-Lite is a Gemini 3-series model for low-latency, high-throughput multimodal processing. ... On B.AI, use the model ID gemini-3.5-flash-lite. ... Context Window: Supports up to 1,048,576 input tokens. Max Output: Supports up to 65,536 output tokens. ... Multimodal Input: Supports text, image, video, audio, and PDF input with text output. Tool Integration: Supports function calling, structured outputs, code execution...