Google announced a set of new models for its Gemini family. The release focuses on refining the balance between performance, efficiency, and cost across different use cases.
Rather than introducing entirely new flagship systems, the updates target the Flash series with incremental but measurable improvements in token usage, speed, and task-specific capabilities.
This approach reflects Google's ongoing effort to make its models more practical for developers and everyday users who need reliable performance without excessive resource demands.
The announcement included supporting images and a short video that illustrated the models in context.
At its core, the update addresses feedback from users and developers who have worked with earlier Flash versions. Common requests centered on reducing verbosity in outputs, improving consistency in multi-step workflows, and lowering overall operational costs.
Google responded with models that maintain or enhance quality while using fewer tokens and executing tasks more directly.
The updates center on improving efficiency, speed, and specialized performance while keeping costs in check.
The models include Gemini 3.6 Flash as the main everyday option, Gemini 3.5 Flash-Lite for high-volume tasks, and Gemini 3.5 Flash Cyber for security work.
Gemini 3.6 Flash builds on the previous 3.5 Flash with changes based on user feedback.
It shows gains in coding, reasoning, and tool use. The model reduces output token usage, reported at around 17% less than its predecessor according to Artificial Analysis data, and up to 65% fewer on certain benchmarks such as DeepSWE.
This efficiency comes with lower pricing at $1.50 per million input tokens and $7.50 per million output tokens. Benchmark improvements include better results on agentic coding tasks, knowledge work evaluations, and computer use scenarios.
The knowledge cutoff date moved forward to March 2026.
Users can access it through the Gemini app by selecting the model in the dropdown, as well as in Google AI Studio, the Gemini API, Android Studio, and enterprise platforms.
One practical example involves extracting natural textures from photos to create design elements or prints for further workflows.
The model handles multimodal inputs including text, images, video, audio, and PDFs within a one million token context window. It supports standard tools such as function calling and computer use.
Gemini 3.5 Flash-Lite targets situations that need high throughput and low latency.
It runs at roughly 350 output tokens per second and carries a lower price point of $0.30 per million input tokens and $2.50 per million output tokens.
The model outperforms earlier Flash-Lite versions on coding, agentic tasks, and long-context work.
It also shows advantages over some prior Flash releases in real-world evaluations.
Availability includes the Gemini app, Google Search, and the same developer and enterprise channels as the other models.
The third addition, Gemini 3.5 Flash Cyber, focuses on cybersecurity.
Built on the 3.5 Flash base, it specializes in identifying and fixing vulnerabilities inside Google's CodeMender code security agent. Multiple instances of the model can work together to generate combined reports.
Google positioned it for lower cost than larger models while reaching competitive performance levels. Access remains limited to trusted partners and governments through a pilot program.
These releases continue Google's pattern of refining its Flash line for practical agentic use cases.
The company noted ongoing testing for Gemini 3.5 Pro and early work on the next major generation. The models rolled out starting the day of the announcement across consumer and developer platforms. Developers and users interested in testing can check the Gemini app or AI Studio for the available options.















































































































































































































































































































































































