Google DeepMind's Gemini 3.6 Flash Cuts Token Costs by Up to 65 Percent
Google DeepMind has released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, reducing token costs for automated engineering workflows by as much as 65 percent.

Google DeepMind has introduced three new models aimed at reducing the cost of running AI agents at scale. The releases, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-focused variant called Gemini 3.5 Flash Cyber, are designed to make autonomous software workflows cheaper and faster to operate.
Gemini 3.6 Flash is the headline model. It is priced at $1.50 per million input tokens and $7.50 per million output tokens. Beyond pricing, it is built to reduce the total number of tokens generated when solving multi-step tasks. Independent testing by Artificial Analysis found that the model cuts output token usage by 17 percent overall compared to its predecessor. On complex coding benchmarks like DeepSWE, token consumption drops by up to 65 percent, while task completion accuracy rises from 37 percent to 49 percent.
Gemini 3.5 Flash-Lite takes a different position. Priced at $0.30 per million input tokens and $2.50 per million output tokens, it is built for throughput, processing 350 output tokens per second. That makes it well suited to high-volume use cases such as search agents and real-time document processing where speed matters more than deep reasoning.
The third model, Gemini 3.5 Flash Cyber, is purpose-built for security work: vulnerability research and patch generation. It will be distributed through Google's CodeMender agent to select partners and government entities, rather than made broadly available through the standard API.
The releases reflect a broader shift in how AI providers compete. As developers build agents that run continuously rather than respond to one-off prompts, inference efficiency matters as much as benchmark scores. Google technical staffer Logan Kilpatrick confirmed that the more capable Gemini 3.5 Pro remains in partner testing, but the current batch of lightweight models is already designed for practical, daily production use. Fewer reasoning loops mean lower infrastructure costs and smaller API bills for the teams building on top of these systems.
For African developers and startups, that distinction is significant. Engineering teams across Lagos, Nairobi, Johannesburg, and Cairo regularly absorb cloud compute bills priced in US dollars while operating in local currencies under sustained exchange rate pressure. A 65 percent reduction in token usage on agent tasks is not a marginal improvement; it is the difference between a product being financially viable and not. Lower costs open the door to building autonomous customer service tools, local-language processing pipelines, and developer utilities that would have been prohibitively expensive at previous pricing levels. Faster inference also helps offset latency challenges that come with geographic distance from primary server infrastructure in Europe and North America.
The real competition in AI infrastructure is no longer about which model scores highest on a benchmark; it is about which one costs the least to run in production.
Source: VentureBeat
Written by
StacyAI-assisted news curation. Every story is reviewed by our editors before publication.


