Gemini 3.5 Flash
Google's agentic workhorse, unveiled at I/O 2026 — Intelligence Index 55, 152 tokens/sec with an 18.7-second first token, 1M-token context, and full multimodal input (text, image, video, audio, PDF) at $1.50/$9 per million tokens.
Model Specs
- Released
- May 2026
- Context window
- 1.0M tokens
- Max output
- 66K tokens
- Capabilities
- reasoningmultimodalfast-inferencefunction-calling
- Modalities
- textvisionaudiovideo
About this model
Gemini 3.5 Flash is Google's mainstream frontier model, announced as generally available at Google I/O on May 19, 2026. It's the default model behind the Gemini app and AI Mode in Search globally, and the engine of Google's new Spark agent — a release that, in Google's framing, bets the next AI wave on agents rather than chatbots. On the Artificial Analysis Intelligence Index it scores 55 (#10 overall), remarkable for a model that streams at 152 tokens/sec with a time-to-first-token of just 18.7 seconds.
The benchmark profile is agent-shaped. Terminal-Bench 2.1: 76.2%. MCP Atlas: 83.6%. OSWorld-Verified (computer-use tasks): 78.4%. SWE-Bench Pro: 55.1%. It also posts 40.2% on Humanity's Last Exam, 72.1% on ARC-AGI-2, and 83.6% on MMMU-Pro for multimodal reasoning. Notably, it beats the bigger Gemini 3.1 Pro on coding and agentic benchmarks while trailing it on the hardest reasoning evals — exactly the trade Google intended for a high-throughput workhorse. Reasoning is tunable across four thinking levels (minimal/low/medium/high, default medium), and the model accepts text, images, video, audio, and PDFs in a 1,048,576-token context window with 65K output tokens.
Pricing is flat — $1.50/M input and $9/M output with no context-length tiering, $0.15/M cached input, and Batch/Flex at half price. That's roughly 3x the old Gemini 3 Flash Preview but about 40% cheaper than Gemini 3.1 Pro, and ~3x cheaper than GPT-5.5. Reach for Gemini 3.5 Flash when you need near-frontier capability at interactive speed: agent loops, multimodal pipelines, high-volume production workloads, and any flow where a 100-second first token would kill the experience.
Key Strengths
Frontier-adjacent intelligence at speed
Intelligence Index 55 (#10 overall) while streaming 152 tokens/sec with an 18.7s first token — an order of magnitude faster to respond than deep-reasoning flagships like GPT-5.5 (~103s) or Claude Fable 5 (~108s).
Agent-first benchmark profile
Terminal-Bench 2.1 76.2%, MCP Atlas 83.6%, OSWorld-Verified 78.4% — built for tool-calling loops and real agent workflows, and it even beats Gemini 3.1 Pro on coding/agentic evals.
Full multimodal input
Text, images, video, audio, and PDFs in one 1M-token context. Analyze a recorded meeting, a slide deck, and a codebase in a single request — most rivals at this speed tier are text+image only.
Flat, predictable pricing
$1.50/M input, $9/M output with no context-length tiering — long-context requests cost the same rate. Batch and Flex tiers halve it; cached input is $0.15/M.
Tunable thinking levels
Four reasoning levels (minimal/low/medium/high) with automatic thought preservation across turns — dial depth up for hard steps and down for fast ones inside the same conversation.
Google-grounded tooling
Native Google Search and Maps grounding, code execution, file search, and URL context — retrieval and verification primitives built into the model rather than bolted on.
How it compares
Gemini 3.5 Flash competes on speed-adjusted intelligence — near-frontier quality at interactive latency and economy pricing.
| vs. Model | Verdict | Outcome |
|---|---|---|
| Gemini 2.0 Flash | A massive generational leap: Intelligence Index 55 vs 19, agentic benchmarks that didn't exist for 2.0 Flash, full thinking-level controls, and stronger multimodality. Gemini 2.0 Flash remains 10x cheaper raw ($0.15/$0.60) for simple high-volume tasks, but 3.5 Flash is the new default for anything requiring real capability. | Wins most cases |
| GPT-5.5 | GPT-5.5 is deeper (Index 60 vs 55, GDPval 84.9%) but ~3x the price ($5/$30 vs $1.50/$9) and ~5x slower to first token (103s vs 18.7s). Flash wins interactive, agentic, and high-volume work; GPT-5.5 wins when maximum reasoning depth justifies the wait and the cost. | Depends |
| Claude Haiku 4.5 | Both target the fast tier, but 3.5 Flash brings substantially more intelligence (Index 55 vs 31), video/audio/PDF input, and Google grounding. Haiku 4.5 counters with sub-second first tokens (0.85s vs 18.7s) and SWE-bench Verified 73.3% for snappy coding flows. Latency-critical UX favors Haiku; capability-per-dollar favors Flash. | Wins most cases |
Pros
- Intelligence Index 55 (#10) at speed-tier latency — 152 t/s, 18.7s first token
- Strong agentic results: Terminal-Bench 76.2%, MCP Atlas 83.6%, OSWorld 78.4%
- Full multimodal input: text, image, video, audio, PDF in 1M context
- Flat $1.50/$9 pricing with no context tiering; Batch/Flex at half price
- Four tunable thinking levels with cross-turn thought preservation
- Native Google Search/Maps grounding and code execution
- Beats Gemini 3.1 Pro on coding/agentic benchmarks at 40% lower price
Things to consider
- Trails deep flagships on hardest reasoning (HLE 40.2% vs Fable 5's 53%)
- 65K output token cap — lower than the 128K of GPT-5.5/Fable 5
- Text-only output: no image/audio generation, no Live API support
- ~3x the price of its Gemini 3 Flash Preview predecessor
- Google recommends dropping temperature/top_p tuning — less sampling control
- January 2025 knowledge cutoff is older than GPT-5.5's December 2025
Best use cases
Production AI agents
The MCP Atlas and Terminal-Bench results plus 18.7s first token make it ideal for agent loops that need many fast tool-calling rounds — exactly what Google's own Spark agent runs on.
Multimodal pipelines
Process video, audio, PDFs, and images in one request — meeting summarization, media QA, document extraction, and visual inspection workflows.
Interactive chat & copilots
Near-frontier quality at latency users will actually wait for. The default-medium thinking level balances depth and responsiveness for assistant UX.
High-volume content generation
At $1.50/$9 flat with 152 t/s throughput, batch summarization, classification, and drafting workloads run ~3x cheaper than GPT-5.5-class models.
Computer-use automation
78.4% on OSWorld-Verified — strong autonomous performance on real desktop tasks for RPA-style automations and UI agents.
Search-grounded research
Built-in Google Search grounding lets the model verify claims and pull fresh information mid-task — useful for research assistants and fact-checked drafting.
How to use it on Renas AI
- 1
Step 1
Open AI Chat on Renas
Navigate to AI Chat in the Renas dashboard and pick the most capable Gemini model available to your plan from the model selector.
- 2
Step 2
Throw multimodal input at it
Don't limit it to text — attach images, documents, or media. The 1M-token multimodal window is the differentiator; use it for context other models can't hold.
- 3
Step 3
Match thinking level to the task
Default medium suits most work. Step up to high for tricky reasoning, drop to minimal/low for classification and extraction where speed matters most.
- 4
Step 4
Iterate at conversational speed
With sub-20-second first tokens, treat it like a real-time collaborator — rapid draft-review-revise loops that would feel glacial on deep-reasoning flagships.
Pricing
Pricing on Renas AI
Pay-as-you-go credits, no API keys, no rate limits.
~500,000 words on a 10,000-credit Spark plan
Frequently asked questions
Other Google models
Other text models on Renas AI
Google's fastest frontier model on Renas
Access leading Gemini models with your Renas AI subscription credits — no API key, no setup, no per-seat fees.
Try Gemini models on Renas