TL;DR
Google now offers 3.6 Flash for better performance and cost efficiency, 3.5 Flash-Lite for speed, and 3.5 Flash Cyber for cybersecurity tasks, all optimized for scaling AI agents.
Key points
- 1
3.6 Flash: Better efficiency and quality: 3.6 Flash improves coding, knowledge work, and multimodal performance over 3.5 Flash. It reduces output tokens by 17% according to the Artificial Analysis Index and achieves up to 65% better performance in benchmarks like DeepSWE. The model also costs less per output token ($1.50/1M input and $7.50/1M output) while maintaining higher precision—showing 49% vs. 37% accuracy in code edits and 83.0% vs. 78.4% in OSWorld-verified tasks. Developers should switch to 3.6 Flash for cost-effective, high-quality agent workflows that handle complex tasks like financial data analysis and report drafting without excessive token usage.
- 2
3.5 Flash-Lite: Fastest for high-volume tasks: 3.5 Flash-Lite delivers 350 output tokens per second and is optimized for low-latency, high-throughput scenarios like agentic search and document processing. It outperforms older versions in benchmarks such as Terminal-Bench 2.1 (54% vs. 31%) and GDM-MRCR v2 (72.2% vs. 60.1%). Customers can configure it to prioritize speed for tasks like receipt translation or e-commerce data synthesis while using higher thinking levels for complex multi-step workflows. Developers should deploy 3.5 Flash-Lite in Google AI Studio or Android Studio for real-time applications where rapid processing is critical, such as generating web design concepts or scaling cybersecurity tasks.
- 3
3.5 Flash Cyber: Cybersecurity-focused model: 3.5 Flash Cyber is a specialized model for finding and fixing vulnerabilities, paired with CodeMender’s agent infrastructure. It’s designed for governments and trusted partners through a limited pilot program to detect security issues faster than traditional systems. The model achieves competitive performance on CyberGym benchmarks and is priced at a lower cost per token than larger models. Organizations should apply for the CodeMender pilot to use this model for proactive vulnerability scanning, especially for critical infrastructure where rapid response to threats is essential.
What changed
Before this update
Developers used older Gemini models that were less efficient for building AI agents at scale.
After this update
Google has introduced 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber to improve token efficiency, reduce latency, and enhance reliability for production AI workflows.
Share this update
This is a summary of an official post from the Google Search Central Blog, provided for quick reading. Google and the Google logo are trademarks of Google LLC; My Tool Studio is not affiliated with Google. Always refer to the original announcement for authoritative guidance.