Close Menu
WithO2WithO2

    Subscribe to Updates

    Get the latest AI News Tools Updates in your Inbox

    What's Hot

    Meta Muse Code Launches: AI Coding Agent 21x Cheaper to Try

    August 6, 2026

    Google DeepMind CEO Steps Down — What It Means for Gemini

    August 6, 2026

    Ode with Anthropic: Inside the $1.5B AI Implementation Firm

    August 6, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    WithO2WithO2
    • AI
    • Blog
    • Business Software
    • Trending News
    • Stories
    WithO2WithO2
    Home » Trending News
    Trending News

    Gemini 3.5 Pro Is Live: 2M Context, Deep Think in Vertex Preview

    By Amitabh SarkarJuly 15, 2026Updated:August 6, 202614 Mins Read12
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Gemini 3.5 Pro launch: Google talent exodus to Anthropic and OpenAI before July 17 launch
    Gemini 3.5 Pro arrives July 17 following the biggest AI talent shuffle of 2026.
    Share
    Facebook Twitter LinkedIn Pinterest Email
    🔴 Update — August 5, 2026 — Leadership shakeup: Google DeepMind confirmed a major restructuring that directly affects Gemini development. Demis Hassabis stepped down as Google DeepMind CEO and becomes Chair of Google DeepMind and Chief Scientist of Alphabet, shifting to long-term AGI strategy. Koray Kavukcuoglu — DeepMind’s CTO for 13 years — was named SVP of Google DeepMind and now owns all Gemini model development, reporting directly to Sundar Pichai. On the same day Kavukcuoglu was announced, two of Gemini’s co-technical leads departed Google: Oriol Vinyals and Quoc Le joined Chief Scientist Jeff Dean (27 years at Google) and Sanjay Ghemawat to co-found Discovery Loop, a public benefit corporation backed by Google as a founding investor. Gemini 3.5 Pro remains unreleased with no confirmed public launch date, now more than 60 days past its original June 2026 target. Alphabet stock fell approximately 5%, erasing roughly $190 billion in market value. Sources: Business Today · The Next Web
    🔵 Update — July 22, 2026: Google released three Flash-tier models on July 21 — Gemini 3.6 Flash ($1.50/$7.50 per M tokens, 17% fewer output tokens vs 3.5 Flash), Gemini 3.5 Flash-Lite ($0.30/$2.50 per M tokens, 350 output tokens/second), and Gemini 3.5 Flash Cyber (governments and trusted partners only, fine-tuned for cybersecurity). Gemini 3.5 Pro remains unreleased — Google said it is “currently testing with partners.” Google also started its most ambitious Gemini 4 pre-training run to date. Details below ↓
    🟡 Update — July 19, 2026: The July 17 Vertex AI preview is now confirmed to be a limited, preview-only rollout — not the full public launch Google originally targeted. As of July 19, public API pricing has not been published and the model is not available outside Vertex preview. The TechTimes (July 16) described this as Google “eyes stopgap release” after missing its third deadline. Polymarket’s prediction market (>23K contracts) currently prices the probability of a full public launch at 81% by July 31, 2026.
    🟢 Update — July 18, 2026: Gemini 3.5 Pro launched on July 17 in limited Vertex preview. Confirmed specs added below. Kimi K3 (2.8T) launched the same day — competitive comparison added.

    Gemini 3.5 Pro launched on July 17, 2026 in a limited Vertex AI preview — matching its widely reported target date after a third delay that had pushed the release from June. Multiple independent trackers confirm the model went live July 17, with a 2.1-million-token context window and Deep Think reasoning mode. Google has not published a standalone launch blog post or confirmed public API pricing as of July 18.

    Status as of July 22, 2026:

    • Launch: ✅ Confirmed live — July 17, 2026, limited Vertex AI preview
    • Context window: 2.1 million tokens (confirmed by multiple independent trackers — no official Google spec sheet yet)
    • Deep Think reasoning: Included; access tier not officially confirmed — reports indicate Ultra plan (pricing unconfirmed)
    • Architecture: Complete rebuild from scratch — not incremental from Gemini 2.5 Pro
    • API pricing: Not publicly confirmed as of July 22
    • GA (General Availability): Not yet — “currently testing with partners,” per Google (July 21)
    • Same-day launch: Moonshot AI’s Kimi K3 (2.8T) also launched July 16–17 at $3/$15 per million tokens

    Table of Contents

    Toggle

    • What Actually Happened on July 17
    • Confirmed: 2M Context Window and Deep Think
    • Kimi K3 Launched the Same Day — What It Means for the Competition
    • What the Gemini 3.5 Pro Leaks Actually Said
    • Why Google Rebuilt the Model From Scratch
    • The Talent Losses That Raised the Stakes
    • Google Ships Three Flash Models on July 21
    • What We Still Don’t Know
    • FAQ

    What Actually Happened on July 17

    Gemini 3.5 Pro went live on July 17 in a limited Vertex AI preview — confirming the widely reported date after a chaotic six-week build cycle that included at least three separate delay announcements. The launch was quiet by Google standards: no standalone blog.google announcement, no public model card, and no official API pricing page as of this writing. Independent model trackers confirmed the July 17 release, with a 2.1-million-token context window placing it second among all tracked frontier models for context depth.

    The limited preview status means general availability is still pending. Enterprise customers with Vertex access can test the model now; broader developer access through Google AI Studio and the public Gemini API has not been announced. This pattern matches how Google rolled out Gemini 3.1 Pro — limited Vertex access first, broader API release weeks later.

    Confirmed: 2M Context Window and Deep Think

    The two headline specs from the pre-launch leaks — the 2-million-token context window and the Deep Think reasoning layer — have been confirmed by independent model trackers, though Google has not yet published an official model card with verified benchmark scores. The context window of 2.1 million tokens (2,097,152 to be exact) is the second-largest of any tracked frontier model. For practical reference, 2 million tokens can hold an entire software codebase, a year of enterprise email, or tens of thousands of rows of structured data — all in a single prompt without chunking.

    Deep Think, the step-by-step reasoning mode, is confirmed present. Access tier and pricing remain unconfirmed — the pre-launch leak placed it on a $250/month Ultra plan; the actual launch tier has not been officially stated. Until Google publishes a pricing page, treat any tier assignment as provisional. What is clear is that Deep Think operates similarly to reasoning modes from OpenAI and Anthropic: trading response speed for deeper problem-solving on multi-step tasks like mathematical proofs, long-horizon code refactors, and policy analysis.

    Kimi K3 Launched the Same Day — What It Means for the Competition

    Moonshot AI launched Kimi K3 on July 16–17, the same window as Gemini 3.5 Pro, and the juxtaposition is striking. Kimi K3 is a 2.8-trillion-parameter open-weight model — the world’s first open 3T-class model — with a 1-million-token context window and native vision. It landed at $3.00 per million input tokens and $15.00 per million output tokens, undercutting Claude Opus 4.8 ($5/$25) and GPT-5.5 ($5/$30) on both sides.

    In blind Arena evaluations, developers chose Kimi K3 ahead of all U.S. frontier models on front-end coding tasks, and independent rankings place it fourth overall — trailing only Claude Fable 5, GPT-5.6 Sol, and closely edging past Claude Opus 4.8. Full open weights are scheduled to drop by July 27, 2026, according to Fortune’s July 17 coverage.

    The comparison with Gemini 3.5 Pro is instructive: Kimi K3 wins on context depth by 1 million tokens (Gemini’s 2.1M vs Kimi’s 1M), Gemini wins on transparency (Kimi K3 is open-weight; Gemini 3.5 Pro is proprietary), and on price Kimi K3 is drastically cheaper if Google’s unconfirmed leaked pricing of ~$12–15/$36–45 per million tokens holds. If you’re a developer choosing between them today, Kimi K3 is available via public API right now; Gemini 3.5 Pro is not. If you’re an enterprise user with Vertex access, Gemini 3.5 Pro’s tighter Google ecosystem integration is the differentiator to test.

    What the Gemini 3.5 Pro Leaks Actually Said

    As reported by TechTimes on July 13, Gemini 3.5 Pro targeted July 17 after a full architectural rebuild. The headline pre-launch leak was the 2-million-token context window, now confirmed by trackers. For comparison, Anthropic’s Claude Fable 5 ships with a 200K standard window, and OpenAI has not publicly confirmed the context size for GPT-5.6 Sol.

    The second major leak concerned Deep Think, a reasoning layer reportedly reserved for the $250-per-month Ultra plan. Now that the model is live, the spec appears confirmed in function — the pricing tier has not been officially stated. Reports also suggested the rebuild specifically targeted three weaknesses: mathematical reasoning, SVG scene generation, and image quality.

    On API pricing, a leaked post circulating on X suggested ~$12–15 per million input tokens and ~$36–45 per million output tokens. These figures remain unconfirmed — Google has still published no official pricing page as of July 22. For context, Gemini 3.1 Pro currently costs $2/$12 per million tokens; the leaked Gemini 3.5 Pro pricing would represent a 6–8× premium, reflecting the compute overhead of Deep Think reasoning.

    Why Google Rebuilt the Model From Scratch

    Gemini 3.5 Pro was originally slated for late June 2026 — making the July 17 launch more than six weeks late. According to StartupFortune, Google delayed after scrapping its base model entirely, citing “quality refinements after early enterprise testing.” A full rebuild this late in a release cycle is unusual — it signals either serious performance problems in the original model or unusually high ambitions for the replacement. Geeky Gadgets reported on July 15 that the model had encountered a third delay, with hallucinations and inconsistent outputs as specific concerns. Despite that, the July 17 limited preview appears to have shipped.

    The Talent Losses That Raised the Stakes

    The delay arrived alongside the most significant talent exodus Google has seen in its AI division. On June 18, Noam Shazeer — Gemini co-lead and co-author of the original Transformer architecture paper — left for OpenAI. The following day, John Jumper, the 2024 Nobel laureate who led AlphaFold, announced his move to Anthropic, along with fellow researchers Jonas Adler and Alexander Pritzel.

    The exits triggered a roughly 7% drop in Alphabet’s stock on June 22, erasing an estimated $270 billion in market value across two trading sessions — a stark signal of how tightly investor confidence in Google is now linked to its AI leadership pipeline.

    The departures matter beyond the stock chart. Shazeer’s institutional knowledge of transformer architecture runs deep, and Jumper’s move to Anthropic directly strengthens the team building Claude. That both men left during the very period when Gemini 3.5 Pro was being rebuilt from scratch adds context the spec sheets won’t include. The limited-preview launch on July 17 is Google’s first public signal that the rebuild survived those losses.

    Google Ships Three Flash Models on July 21 — Gemini 3.5 Pro Still Absent

    Google released three Flash-tier models on July 21, 2026 — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber — while Gemini 3.5 Pro remained unreleased five weeks past its July 17 limited preview. All three sit in Google’s Flash tier, tuned for speed, cost, and high-volume agentic work rather than maximum reasoning depth.

    Gemini 3.6 Flash is the primary workhorse update. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it costs less than 3.5 Flash while generating 17% fewer output tokens on the Artificial Analysis Index — meaning lower total cost per agentic task. On coding benchmarks, 3.6 Flash scores 49% on DeepSWE versus 37% for 3.5 Flash, and improves ML research performance (MLE Bench: 63.9% vs. 49.7%). Computer use is now a built-in client-side tool via the Gemini API.

    Gemini 3.5 Flash-Lite targets high-throughput agentic workloads. At $0.30 per million input tokens and $2.50 per million output tokens, it is Google’s cheapest 3.5-class model, running at 350 output tokens per second per Artificial Analysis. On coding benchmarks it outperforms Gemini 3 Flash — 54.2% vs. 49.6% on SWE-Bench Pro and 74.0% vs. 65.1% on OSWorld-Verified — making it a faster and cheaper option for workloads previously routed to Gemini 3 Flash.

    Gemini 3.5 Flash Cyber, a fine-tuned variant of 3.5 Flash optimized for finding and fixing cybersecurity vulnerabilities, will be available exclusively to governments and trusted partners via Google’s CodeMender platform as part of a limited-access pilot. Google did not publish general API pricing for Cyber; the model is not available to individual developers.

    On Gemini 3.5 Pro, Google’s official blog stated the model is “currently testing with partners” and the company plans to “make it broadly available as soon as it’s ready.” Bloomberg reported meeting internal performance targets has been a challenge — consistent with the model’s history of missed deadlines since its original late-June target.

    Google also announced it has begun “the most ambitious pre-training run yet” for Gemini 4 — the first public acknowledgment of the next-generation model’s training timeline.

    What We Still Don’t Know

    Pricing remains the biggest open question. Google has published no official API pricing page for Gemini 3.5 Pro as of July 22. The leaked $12–15/$36–45 per million token figure may or may not hold when general availability is announced. General availability itself has no announced date. No independent benchmark scores have been verified with a source URL — third-party ranking sites report the model is live but show no confirmed leaderboard position. We will update this article when Google publishes an official model card, pricing page, or GA announcement.

    The timing follows a crowded frontier-model summer: OpenAI’s GPT-5.6 Sol, Terra, and Luna family entered limited preview on June 26, Moonshot AI’s Kimi K3 launched July 16–17, and Google’s own Gemini 3.5 Flash has been serving as the family’s speed tier since earlier this year.

    💡 Our Take: Google’s July 21 Flash releases are a pragmatic move — ship what’s ready, buy time on 3.5 Pro. Gemini 3.6 Flash’s 17% token-efficiency gain and lower price are real, developer-facing improvements. But the bigger story is what’s missing: 3.5 Pro, announced at Google I/O months ago, is still “testing with partners” while Bloomberg flags internal performance challenges. The Gemini 4 pre-training announcement reads as a hedge — reminding the market that Google is looking ahead even as the current Pro-tier release remains indefinitely delayed.

    FAQ

    Did Gemini 3.5 Pro actually launch on July 17?

    Yes. Gemini 3.5 Pro launched on July 17, 2026 in a limited Vertex AI preview. It is not yet generally available — broader developer access through Google AI Studio and the public Gemini API has not been announced as of July 22.

    What is the context window for Gemini 3.5 Pro?

    Independent model trackers confirm a context window of 2.1 million tokens (2,097,152), making it the second-largest of any tracked frontier model. Google has not published an official model card with this figure verified directly.

    How much will Gemini 3.5 Pro cost via API?

    Google has not published official pricing as of July 22, 2026. A pre-launch leak suggested ~$12–15 per million input tokens and ~$36–45 per million output tokens — treat this as unconfirmed until Google publishes a pricing page. For comparison, Kimi K3 launched the same day at $3/$15 per million tokens, and Gemini 3.6 Flash (released July 21) costs $1.50/$7.50 per million tokens.

    Does Gemini 3.5 Pro have Deep Think reasoning?

    Yes. Deep Think reasoning mode is confirmed present in the limited preview. The specific subscription tier required to access it has not been officially confirmed — pre-launch leaks suggested the $250/month Ultra plan, but this has not been verified by Google.

    How does Gemini 3.5 Pro compare to Kimi K3?

    Gemini 3.5 Pro has a larger context window (2.1M vs 1M tokens) and tighter Google ecosystem integration. Kimi K3 is available via public API right now, is open-weight (full weights dropping by July 27), and is significantly cheaper — $3/$15 per million tokens vs Gemini 3.5 Pro’s unconfirmed leaked pricing of $12–15/$36–45. For developers, Kimi K3 is the accessible option today; Gemini 3.5 Pro requires Vertex enterprise access.

    Why did Google rebuild Gemini 3.5 Pro from scratch?

    Google reportedly scrapped the original base model and performed a full architectural rebuild, citing quality refinements after early enterprise testing. The rebuild ran more than six weeks past the original late-June target and reportedly targeted math reasoning, SVG scene generation, and image quality improvements.

    Who left Google’s AI team before the launch?

    Noam Shazeer, Gemini co-lead and Transformer co-author, left for OpenAI on June 18, 2026. John Jumper, 2024 Nobel laureate and AlphaFold lead, moved to Anthropic on June 19. The exits erased roughly $270 billion in Alphabet market cap across two trading sessions.

    What are Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber?

    Google released all three models on July 21, 2026. Gemini 3.6 Flash costs $1.50/$7.50 per million tokens and uses 17% fewer output tokens than 3.5 Flash. Gemini 3.5 Flash-Lite costs $0.30/$2.50 per million tokens and runs at 350 output tokens per second — optimized for high-throughput agentic workloads. Gemini 3.5 Flash Cyber is a security-specialized variant available only to governments and trusted partners via Google’s CodeMender platform; it is not available via the public API.

    What is Gemini 4?

    On July 21, 2026, Google confirmed it has begun its “most ambitious pre-training run yet” for Gemini 4. No further details — architecture, timeline, or capabilities — have been disclosed. This is the first official acknowledgment that Gemini 4 is in active training.

    Updated: August 5, 2026

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Amitabh Sarkar
    • Website

    I am a software engineer, I have a passion for working with cutting-edge technologies and staying up-to-date with the latest developments in the field. In my articles, I share my knowledge and insights on a range of topics, including business software, how to set up tools, and the latest trends in the tech industry.

    Related Posts

    Meta Muse Code Launches: AI Coding Agent 21x Cheaper to Try

    August 6, 2026

    Google DeepMind CEO Steps Down — What It Means for Gemini

    August 6, 2026

    Ode with Anthropic: Inside the $1.5B AI Implementation Firm

    August 6, 2026

    Comments are closed.

    Don't Miss
    Trending News

    Meta Muse Code Launches: AI Coding Agent 21x Cheaper to Try

    By Amitabh SarkarAugust 6, 2026

    Published August 6, 2026 Meta launched Muse Code on August 5, 2026 — its first…

    Google DeepMind CEO Steps Down — What It Means for Gemini

    August 6, 2026

    Ode with Anthropic: Inside the $1.5B AI Implementation Firm

    August 6, 2026

    Microsoft Copilot Studio’s AI Workflow Designer Is Now Live for All

    August 6, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    Our Picks

    12 Best Project Management Software Tools in 2026

    August 1, 2026

    9 Best Ecommerce Platforms Compared (2026)

    July 30, 2026

    Rippling vs Gusto vs BambooHR: Full HRMS Comparison 2026

    July 15, 2026

    Best Ecommerce Platform 2026: Top 10 Options Compared

    July 5, 2026
    Editors Picks

    Meta Muse Code Launches: AI Coding Agent 21x Cheaper to Try

    August 6, 2026

    Google DeepMind CEO Steps Down — What It Means for Gemini

    August 6, 2026

    Ode with Anthropic: Inside the $1.5B AI Implementation Firm

    August 6, 2026

    Microsoft Copilot Studio’s AI Workflow Designer Is Now Live for All

    August 6, 2026
    About Us
    About Us

    Your Source for Innovation: Discover in-depth guides, solutions, and tools tailored to modern business challenges.

    Links
    • Blog
    • Privacy Policy
    • Contact WithO2.com
    • Terms and Conditions
    Facebook X (Twitter) Instagram Pinterest
    • About
    • Editorial Policy
    • Contact
    • Privacy Policy
    • Terms
    © 2026 WITHO2.COM

    Type above and press Enter to search. Press Esc to cancel.