Gemini

Gemini 3.5 Flash

A frontier-level ultra-fast native five-modal large model developed by Google DeepMind, equipped with dynamic multi-level deep thinking, stable agent orchestration, flagship coding capability and million-token long context, delivering frontier intelligence at half the cost, suitable for enterprise-scale agent clusters, full-stack R&D, bulk multimedia processing, massive document governance and high-concurrency commercial API scenarios.

Modalities
Chat
Starting price
From $6.75 / 1M out
Context
1.0M context

Google

README

Supported Functionality

ItemSpecification
InputText, Image, Video, Audio, PDF
OutputText
Context1,048,576 tokens
Max Output65,536 tokens
Vision✓ Supported
Function Calling✓ Supported

Description

Gemini 3.5 Flash is a stable, generally available model released by Google on May 19, 2026. Built on the Gemini 3 Flash reasoning foundation, it uses a natively multimodal reasoning design. At launch, it was Google's most intelligent Flash model for agentic execution, coding, and long-horizon tasks at scale. It remains a stable endpoint with no announced shutdown date, although it is no longer the latest Flash version. Its knowledge cutoff is January 2025, and Google has not disclosed its parameter count.

It combines a 1,048,576-token input context, a 65,536-token maximum output, and Flash-class speed with four thinking levels, cross-turn thought preservation, combined tool use, and Computer Use in preview. Google's May 2026 model card reports scores of 76.2% on Terminal-Bench 2.1, 83.6% on MCP Atlas, 78.4% on OSWorld-Verified, and 83.6% on MMMU-Pro.

Key Capabilities

  • Configurable Deep Reasoning: Four thinking levels—minimal, low, medium, and high—provide flexible control over speed, cost, and reasoning depth.
  • Agentic Execution: Supports subagent deployment, multi-step problem solving, and long-horizon tool use, with an official MCP Atlas score of 83.6%.
  • Coding and Iterative Development: Enables rapid exploration, code generation, execution, and refactoring, scoring 76.2% on Terminal-Bench 2.1 and 55.1% on SWE-Bench Pro Public.
  • Computer Use: Built-in Computer Use in preview can observe, reason, and act across browser, mobile, and desktop environments, scoring 78.4% on OSWorld-Verified.
  • Native Multimodal Understanding: Jointly processes text, images, video, audio, and PDFs, with scores of 83.6% on MMMU-Pro and 84.2% on CharXiv.
  • Million-Token Long Context: A 1,048,576-token input window can hold large codebases, long videos, books, and multi-document collections for long-horizon analysis.
  • Grounding and Combined Tools: Google Search, Google Maps grounding, File Search, URL Context, Code Execution, and custom functions can cooperate in one request.

Technical Strengths

FeatureBenefit
Native Multimodal ReasoningUnderstands text, images, video, audio, and PDFs in one context, reducing cross-model orchestration and information-conversion loss.
Thought PreservationAutomatically carries intermediate reasoning context across turns, improving continuity in iterative debugging, refactoring, and multi-step tasks.
Four Thinking Levelsminimal through high provide clear quality, latency, and cost controls, with medium as the balanced default.
Combined Tool UseMixes search, URL context, code execution, and custom functions within one request, shortening agent workflows.
Built-In Computer UseEnables visual agents across browser, mobile, and desktop environments without requiring a separate computer-use model.
Flexible Inference and CachingBatch, Flex, Priority inference, and Context Caching support deployment choices based on scale, latency, and cost.

Pricing

Token-based pricing

Our pricing is based on image and text token usage. The final cost depends on the tokens consumed.

Token TypeLinkAI PriceOfficial Price
Input$1.125 / 1M tokens$1.5 / 1M tokens
Cached input$0.1125 / 1M tokens$0.15 / 1M tokens
Output$6.75 / 1M tokens$9 / 1M tokens
Reasoning output$6.75 / 1M tokens$9 / 1M tokens

More from Google