Renas AI Logo
Google
Agentic Workhorseby Google

Gemini 3.5 Flash

Google's agentic workhorse, unveiled at I/O 2026 — Intelligence Index 55, 152 tokens/sec with an 18.7-second first token, 1M-token context, and full multimodal input (text, image, video, audio, PDF) at $1.50/$9 per million tokens.

Model Specs

Released
May 2026
Context window
1.0M tokens
Max output
66K tokens
Capabilities
reasoningmultimodalfast-inferencefunction-calling
Modalities
textvisionaudiovideo
Intelligence
55/ 100
#6 of 17 in category
Output speed
152.2t/s
#2 of 14 in category
Renas credits
0.02/ word
#4 of 17 in category
Knowledge cutoff
Jan 2025

About this model

Gemini 3.5 Flash is Google's mainstream frontier model, announced as generally available at Google I/O on May 19, 2026. It's the default model behind the Gemini app and AI Mode in Search globally, and the engine of Google's new Spark agent — a release that, in Google's framing, bets the next AI wave on agents rather than chatbots. On the Artificial Analysis Intelligence Index it scores 55 (#10 overall), remarkable for a model that streams at 152 tokens/sec with a time-to-first-token of just 18.7 seconds.

The benchmark profile is agent-shaped. Terminal-Bench 2.1: 76.2%. MCP Atlas: 83.6%. OSWorld-Verified (computer-use tasks): 78.4%. SWE-Bench Pro: 55.1%. It also posts 40.2% on Humanity's Last Exam, 72.1% on ARC-AGI-2, and 83.6% on MMMU-Pro for multimodal reasoning. Notably, it beats the bigger Gemini 3.1 Pro on coding and agentic benchmarks while trailing it on the hardest reasoning evals — exactly the trade Google intended for a high-throughput workhorse. Reasoning is tunable across four thinking levels (minimal/low/medium/high, default medium), and the model accepts text, images, video, audio, and PDFs in a 1,048,576-token context window with 65K output tokens.

Pricing is flat — $1.50/M input and $9/M output with no context-length tiering, $0.15/M cached input, and Batch/Flex at half price. That's roughly 3x the old Gemini 3 Flash Preview but about 40% cheaper than Gemini 3.1 Pro, and ~3x cheaper than GPT-5.5. Reach for Gemini 3.5 Flash when you need near-frontier capability at interactive speed: agent loops, multimodal pipelines, high-volume production workloads, and any flow where a 100-second first token would kill the experience.

Key Strengths

Frontier-adjacent intelligence at speed

Intelligence Index 55 (#10 overall) while streaming 152 tokens/sec with an 18.7s first token — an order of magnitude faster to respond than deep-reasoning flagships like GPT-5.5 (~103s) or Claude Fable 5 (~108s).

Agent-first benchmark profile

Terminal-Bench 2.1 76.2%, MCP Atlas 83.6%, OSWorld-Verified 78.4% — built for tool-calling loops and real agent workflows, and it even beats Gemini 3.1 Pro on coding/agentic evals.

Full multimodal input

Text, images, video, audio, and PDFs in one 1M-token context. Analyze a recorded meeting, a slide deck, and a codebase in a single request — most rivals at this speed tier are text+image only.

Flat, predictable pricing

$1.50/M input, $9/M output with no context-length tiering — long-context requests cost the same rate. Batch and Flex tiers halve it; cached input is $0.15/M.

Tunable thinking levels

Four reasoning levels (minimal/low/medium/high) with automatic thought preservation across turns — dial depth up for hard steps and down for fast ones inside the same conversation.

Google-grounded tooling

Native Google Search and Maps grounding, code execution, file search, and URL context — retrieval and verification primitives built into the model rather than bolted on.

How it compares

Gemini 3.5 Flash competes on speed-adjusted intelligence — near-frontier quality at interactive latency and economy pricing.

vs. ModelVerdictOutcome
GPT-5.5GPT-5.5 is deeper (Index 60 vs 55, GDPval 84.9%) but ~3x the price ($5/$30 vs $1.50/$9) and ~5x slower to first token (103s vs 18.7s). Flash wins interactive, agentic, and high-volume work; GPT-5.5 wins when maximum reasoning depth justifies the wait and the cost.Depends
Claude Haiku 4.5Both target the fast tier, but 3.5 Flash brings substantially more intelligence (Index 55 vs 31), video/audio/PDF input, and Google grounding. Haiku 4.5 counters with sub-second first tokens (0.85s vs 18.7s) and SWE-bench Verified 73.3% for snappy coding flows. Latency-critical UX favors Haiku; capability-per-dollar favors Flash.Wins most cases

Pros

  • Intelligence Index 55 (#10) at speed-tier latency — 152 t/s, 18.7s first token
  • Strong agentic results: Terminal-Bench 76.2%, MCP Atlas 83.6%, OSWorld 78.4%
  • Full multimodal input: text, image, video, audio, PDF in 1M context
  • Flat $1.50/$9 pricing with no context tiering; Batch/Flex at half price
  • Four tunable thinking levels with cross-turn thought preservation
  • Native Google Search/Maps grounding and code execution
  • Beats Gemini 3.1 Pro on coding/agentic benchmarks at 40% lower price

Things to consider

  • Trails deep flagships on hardest reasoning (HLE 40.2% vs Fable 5's 53%)
  • 65K output token cap — lower than the 128K of GPT-5.5/Fable 5
  • Text-only output: no image/audio generation, no Live API support
  • ~3x the price of its Gemini 3 Flash Preview predecessor
  • Google recommends dropping temperature/top_p tuning — less sampling control
  • January 2025 knowledge cutoff is older than GPT-5.5's December 2025

Best use cases

Production AI agents

The MCP Atlas and Terminal-Bench results plus 18.7s first token make it ideal for agent loops that need many fast tool-calling rounds — exactly what Google's own Spark agent runs on.

Multimodal pipelines

Process video, audio, PDFs, and images in one request — meeting summarization, media QA, document extraction, and visual inspection workflows.

Interactive chat & copilots

Near-frontier quality at latency users will actually wait for. The default-medium thinking level balances depth and responsiveness for assistant UX.

High-volume content generation

At $1.50/$9 flat with 152 t/s throughput, batch summarization, classification, and drafting workloads run ~3x cheaper than GPT-5.5-class models.

Computer-use automation

78.4% on OSWorld-Verified — strong autonomous performance on real desktop tasks for RPA-style automations and UI agents.

Search-grounded research

Built-in Google Search grounding lets the model verify claims and pull fresh information mid-task — useful for research assistants and fact-checked drafting.

How to use it on Renas AI

  1. 1

    Step 1

    Open AI Chat on Renas

    Navigate to AI Chat in the Renas dashboard and pick the most capable Gemini model available to your plan from the model selector.

  2. 2

    Step 2

    Throw multimodal input at it

    Don't limit it to text — attach images, documents, or media. The 1M-token multimodal window is the differentiator; use it for context other models can't hold.

  3. 3

    Step 3

    Match thinking level to the task

    Default medium suits most work. Step up to high for tricky reasoning, drop to minimal/low for classification and extraction where speed matters most.

  4. 4

    Step 4

    Iterate at conversational speed

    With sub-20-second first tokens, treat it like a real-time collaborator — rapid draft-review-revise loops that would feel glacial on deep-reasoning flagships.

Pricing

Pricing on Renas AI

Pay-as-you-go credits, no API keys, no rate limits.

0.02credits per word

~500,000 words on a 10,000-credit Spark plan

Included in every paid plan
No separate API key or setup
Predictable per-word credit cost
Commercial use rights for all output

Frequently asked questions

Google's fastest frontier model on Renas

Access leading Gemini models with your Renas AI subscription credits — no API key, no setup, no per-seat fees.

Try Gemini models on Renas
Gemini 3.5 Flash — Google's Fast Agentic Model | Renas AI | Renas AI