The AI update that makes running enterprise agents cheaper

Gemini 3.5 Flash-Lite runs at 350 output tokens per second, making it Google's fastest model in this specific weight class.

Navi Mumbai | editorial@unboxdailyhq.com
At Unbox Daily HQ, discovery matters more than speed. If it's here, we believe it's worth your time.

The Essentials

  • Google expands its artificial intelligence lineup with three new models aimed directly at developer efficiency and agentic workflows.
  • Gemini 3.6 Flash prices input tokens at $1.50 per million while introducing built-in computer use tools for enterprise automation.
  • Developers building automated applications can now process multi-step tasks faster while spending significantly less on API calls.

The Pulse

Google‘s release of Gemini 3.6 Flash, 3.5 Flash-Lite, and the restricted 3.5 Flash Cyber directly targets the growing developer demand for efficient agent orchestration. Rather than just pushing for larger parameter counts, this update optimises token efficiency and latency for production environments. The 3.6 Flash variant emerges as the primary workhorse, reducing output token consumption by 17 per cent compared to its predecessor on the Artificial Analysis Index.

For Indian developers and businesses integrating artificial intelligence, this translates to reduced operational overhead when scaling applications. The models are available globally and in India starting today through Google AI Studio and the Gemini application.

The introduction of the Cyber variant signals a bifurcation in Google’s deployment strategy. By restricting 3.5 Flash Cyber entirely to governments and vetted partners through the CodeMender platform, the company acknowledges the dual-use risks of frontier code-security tools. Meanwhile, the consumer-facing Flash-Lite model is already rolling out across Google Search, quietly replacing older infrastructure for high-throughput queries.

The Breakdown

Shop NowAD

Sponsored: Unbox Daily HQ earns a commission if you buy through these links, at no extra cost to you. Prices shown are subject to change, and the actual price on Amazon at the time of purchase may vary from what is displayed here.

Under the hood, Gemini 3.6 Flash achieves a 49 per cent success rate on the DeepSWE benchmark, drastically reducing unwanted code edits and execution loops compared to the 37 per cent scored by 3.5 Flash. The 3.5 Flash-Lite model delivers a throughput of 350 output tokens per second, with configurable thinking levels that allow developers to dictate compute intensity based on the workload. Both models now feature upgraded safeguards against chemical, biological, radiological, and nuclear misuse, actively resisting jailbreak attempts while minimising false refusals for legitimate queries. On the multimodal front, 3.6 Flash registers a score of 1421 on GDPval-AA v2 for knowledge work, handling document parsing and chart analysis with measurable precision.

The Distinction

The structural separator for Gemini 3.6 Flash and 3.5 Flash-Lite is the native integration of computer use as a built-in client-side tool. While competing application programming interfaces require developers to build intermediary scaffolding to let an artificial intelligence operate a desktop or browser, these new models handle graphical interfaces and agentic workflows directly out of the box. This architectural shift from a pure text processor to a direct system operator significantly reduces the reasoning steps and separate tool calls required to execute complex, multi-step tasks.

The Snapshot

SpecificationDetails
Model VariantsGemini 3.6 Flash, 3.5 Flash-Lite, 3.5 Flash Cyber
3.6 Flash Input Pricing$1.50 (approx. ₹125) per 1M tokens
3.6 Flash Output Pricing$7.50 (approx. ₹625) per 1M tokens
3.5 Flash-Lite Input Pricing$0.30 (approx. ₹25) per 1M tokens
3.5 Flash-Lite Output Pricing$2.50 (approx. ₹210) per 1M tokens
Flash-Lite Throughput350 output tokens per second
Key UpgradesBuilt-in computer use tool, configurable thinking levels
Cyber Variant AvailabilityRestricted to governments/trusted partners via CodeMender
India AvailabilityAvailable today via Google AI Studio and Gemini app

The Big Picture

The artificial intelligence landscape is shifting from a race for the smartest model to a race for the most economically viable agent. While OpenAI and Anthropic push smaller, faster models like GPT-4o mini and Claude 3.5 Haiku, Google’s aggressive pricing on the Flash series aims directly at high-volume enterprise users. In the Indian tech sector, where startup margins are heavily dictated by cloud compute and API costs, dropping the input price of a highly capable model to under a dollar per million tokens forces competitors to rethink their own billing structures.

The India Prospective

For the Indian developer community and tech startups in Bengaluru and Pune, API pricing dictates product viability. Accessing 3.5 Flash-Lite at approximately ₹25 per million input tokens allows local founders to build high-volume automated workflows, such as customer support agents and document processors, without burning through runway funding. Native integration within Android Studio also directly benefits India’s massive mobile development workforce, enabling faster local deployment of intelligent features.

The Inside Intel

Despite being branded as a lightweight option, the 3.5 Flash-Lite model actually outperforms the heavier Gemini 3 Flash on the SWE-Bench Pro evaluations, scoring 54.2 per cent against the older model’s 49.6 per cent. It is incredibly rare for a lower-tier efficiency model to beat its immediate premium predecessor in complex coding benchmarks, highlighting how rapidly Google is refining its training architecture.

The Unboxed Truth

Unbox Daily HQ considers this the most practical API update for software builders this quarter, not because of raw intelligence gains, but because Google has drastically lowered the financial barrier to running constant, multi-step background tasks.

This update is specifically for a 32-year-old software engineer or technical founder in Bengaluru who is currently throttling their application’s features to keep monthly server costs down. At approximately ₹210 per million output tokens for Flash-Lite, running a dedicated data-parsing agent now costs less than a daily cup of artisanal coffee. That is honest value for a tool that can process thousands of user queries per minute.

The single factor that makes these models worth deploying over competing cheap APIs is the native computer use tool. By removing the need to write custom code just to help the AI interact with a screen or browser, Google is giving developers back hours of their week.

Best for: A technical founder in Bengaluru who needs to run high-volume automated data processing without inflating their monthly cloud bill.

Who Is This For: Perfect for 28 to 42-year-old software developers and startup operators in India who require fast, reliable AI agents for enterprise applications.

Google AI Studio

The Source

Google

The Query

How much does Gemini 3.6 Flash cost in India?

Gemini 3.6 Flash costs $1.50 (approximately ₹125) per million input tokens and $7.50 (approximately ₹625) per million output tokens. It is available in India today through Google AI Studio and the Gemini application. The model reduces output token usage by 17 per cent compared to 3.5 Flash.

How does Gemini 3.5 Flash-Lite differ from Gemini 3 Flash?

Gemini 3.5 Flash-Lite outperforms Gemini 3 Flash on coding evaluations like SWE-Bench Pro, scoring 54.2 per cent against 49.6 per cent. It delivers a high throughput of 350 output tokens per second while offering built-in computer use capabilities. It costs significantly less at $0.30 per million input tokens.

Is Gemini 3.5 Flash-Lite worth using for software developers in India?

Gemini 3.5 Flash-Lite is honest value for technical founders and developers in Bengaluru running high-volume automated workflows. At approximately ₹210 per million output tokens, running background data-parsing agents costs less than a daily cup of coffee. The native computer use tool eliminates custom scaffolding, saving significant technical effort.

Headshot of Ashfaque, an udhq social strategist with dark hair and a maroon shirt, smiling against a plain white background.
Ashfaque S.

With 15+ years across technology infrastructure and digital ecosystems, Ashfaque brings rigorous systems thinking to every story he covers. At Unbox Daily HQ, he researches, tests, and evaluates launches across Technology, Health, Sports, and Business, interrogating claims against real-world Indian conditions before a single word is published. His editorial standard is simple: verified first, published second.
For editorial queries, launch coverage requests, or collaborations, reach out to Ashfaque S. directly at ashfaques@unboxdailyhq.com