Gemini

Gemini 3 Flash Preview

A cost-efficient ultra-fast multimodal large model developed by Google DeepMind, equipped with dynamically adjustable deep reasoning, native four-modal parsing, million-token long context and stable batch tool calling, optimized for low-latency high-concurrency scenarios including enterprise agents, high-frequency API services, coding assistance, bulk audio-video processing and consumer chat applications。

Modalities
Chat
Starting price
From $2.25 / 1M out
Context
100K context

Google

README

Supported Functionality

ItemSpecification
InputText, Image, Video, Audio, PDF
OutputText
Context1,048,576 tokens
Max Output65,536 tokens
Vision✓ Supported
Function Calling✓ Supported

Description

Gemini 3 Flash Preview is Google's natively multimodal reasoning model, released on December 17, 2025. Built on the Gemini 3 Pro reasoning foundation, it delivers advanced reasoning, agentic, and vibe-coding capabilities at Flash-level latency. Its knowledge cutoff is January 2025, and Google has not disclosed its parameter count. As of August 26, 2026, no shutdown date has been announced for the preview endpoint, but Google recommends migrating to the stable gemini-3.6-flash.

It supports a 1,048,576-token input context, a 65,536-token maximum output, four thinking levels, Structured Outputs, and fully multimodal input. Google's launch results report 90.4% on GPQA Diamond, 33.7% on Humanity's Last Exam without tools, 81.2% on MMMU-Pro, and 78% on SWE-bench Verified. Google also states that it is three times faster than Gemini 2.5 Pro and uses 30% fewer tokens on average over typical traffic.

Key Capabilities

  • Configurable Advanced Reasoning: Supports minimal, low, medium, and high thinking levels, with high as the default, to balance quality, latency, and cost by task complexity.
  • Agentic Coding: Targets high-frequency iterative development, tool use, and interactive building, with an official SWE-bench Verified score of 78%.
  • Native Multimodal Understanding: Jointly analyzes text, images, video, audio, and PDFs, with an official MMMU-Pro score of 81.2%.
  • Complex Knowledge Reasoning: Scores 90.4% on GPQA Diamond and 33.7% on Humanity's Last Exam without tools for advanced analysis and planning.
  • Million-Token Long Context: A 1,048,576-token input window can hold large codebases, long videos, books, and multi-document collections, with up to 65,536 output tokens.
  • Agentic Vision: Plans zooming, inspection, and code-based image processing to ground visual answers in more specific evidence.
  • Grounding and Combined Tools: Supports Google Search, Google Maps grounding, File Search, URL Context, Code Execution, Computer Use, and custom functions.

Technical Strengths

FeatureBenefit
Pro Reasoning Foundation with Flash LatencyCombines advanced reasoning with fast responses for interactive coding and agent applications.
Four Thinking Levelsminimal through high let one model cover fast question answering and difficult analysis with flexible reasoning depth.
Native Multimodal ArchitectureProcesses text, images, video, audio, and PDFs in one context, reducing cross-model transformation and orchestration costs.
Agentic Vision and Code ExecutionActively crops, zooms, and processes visual content for more verifiable analysis of charts, interfaces, and detailed images.
Combined Tools and Structured OutputsBuilt-in tools and custom functions can cooperate within one workflow, while structured output supports reliable downstream integration.
Multiple Inference Consumption ModesContext Caching, Batch, Flex, and Priority inference provide deployment choices for throughput, latency, and cost.

Pricing

Token-based pricing

Our pricing is based on image and text token usage. The final cost depends on the tokens consumed.

Token TypeLinkAI PriceOfficial Price
Input$0.375 / 1M tokens$0.5 / 1M tokens
Cached input$0.0375 / 1M tokens$0.05 / 1M tokens
Output$2.25 / 1M tokens$3 / 1M tokens
Reasoning output$2.25 / 1M tokens$3 / 1M tokens

More from Google