Google launched Gemini 3.8 Flash on 2 Sept 2026 — the fourth Flash-tier model in four months — pricing it at $0.75 per million input tokens and setting it as the default model on the Gemini Managed Agents platform for enterprise agentic workflows.

What Gemini 3.8 Flash Does

Gemini 3.8 Flash is a reasoning model designed for long-horizon agentic workflows, including multi-step legal document review, financial analysis, and autonomous coding. Google describes the model as “our most intelligent Flash model.” It outperforms Gemini 3.7 Flash on two enterprise-specific benchmarks: Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark. On DeepSWE v1.1 — a long-horizon software engineering evaluation — it outperforms several larger frontier models, according to Artificial Analysis.

The model supports a 1-million-token context window and a maximum output of 64,000 tokens per call — sufficient to generate complete contracts, filings, or financial reports without truncation. Developers can select from 3 reasoning depths — low, medium, or high — to tune speed against thoroughness per task.

Pricing and Availability

Google has set an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens, valid through December 31, 2026. That makes Gemini 3.8 Flash 6 to 7 times cheaper than Anthropic’s Claude Opus 5, and cheaper than OpenAI’s GPT-5.6 Sol.

The model is available through Google AI Studio and the Gemini API. It is also integrated into the Gemini app (Pro and Ultra tiers), AI Mode, and Gemini in Google Sheets. Google released a companion model alongside the main launch: Gemini 3.8 Flash Cyber, designed for cybersecurity agentic workflows and available to commercial customers.

Why Business Teams Are Paying Attention

For teams running multi-step AI agents for enterprise tasks — CRM automation, financial analysis, compliance review — cost-per-token is the primary constraint on deployment scale. A workflow running against Claude Opus 5 at comparable input volume costs 6 to 7 times more per million tokens. Gemini 3.8 Flash’s 64,000-token maximum output means the model handles document-length outputs — contracts, regulatory filings, multi-section reports — in a single API call.

Gemini 3.8 Flash became the default model on Google’s Gemini Managed Agents platform at launch, according to Google’s enterprise documentation. The Managed Agents platform is what enterprise customers use to build and deploy autonomous multi-step agents without writing scheduling or orchestration infrastructure.

The tunable thinking levels — low, medium, high — give operations teams a cost dial: routine triage tasks can run at low reasoning depth, while high-stakes document analysis runs at high. That flexibility is absent in most fixed-inference-cost APIs.

Competitive Landscape

Google has released 4 Flash-tier models in under four months: Gemini 3.5, 3.6, 3.7, and 3.8 Flash. Each has been priced below OpenAI and Anthropic equivalents while extending agent benchmark coverage. The Register characterized the 3.8 Flash launch as “Google reminds everyone it’s still in the race” — reflecting a sustained market-positioning strategy rather than a single-category breakthrough.

Anthropic’s Claude Fable 5.1 remains the default reasoning model in Salesforce Claudeforce and Agentforce. OpenAI’s GPT-5.6 Sol sits at a higher price point. Gemini 3.8 Flash’s $0.75/M input rate creates direct cost pressure on both at the agentic workload tier.

Teams evaluating the best AI tools for business now have benchmark-backed coverage for legal and financial agent use cases from Google at an introductory price point through the end of 2026.

Our Take

Google is running a volume strategy — 4 Flash models in 4 months — sustaining cost pressure on rivals while enterprise customers consolidate agent stacks. At $0.75 per million input tokens, Gemini 3.8 Flash is the cheapest credible reasoning model for multi-step business workflows available today. Teams that deferred agentic deployments on cost grounds now have a concrete benchmark-backed data point to evaluate against.


For Context — witho2.com’s coverage of Google’s agentic AI model track:

Share.

I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

Comments are closed.

Exit mobile version