Gemini

Gemini 2.5 Flash

Google's next-gen lightweight multimodal model featuring ultra-fast inference, native multimodal fusion, and massive context capacity, optimized for high-frequency interactions, real-time responses, and cost-efficient data processing.

Modalities
Chat
Starting price
From $0.0225 / call
Context
100K context

Google

README

Supported Functionality

ItemSpecification
InputText, image, video, audio
OutputText
Context1,048,576 tokens
Max Output65,536 tokens
Vision✓ Supported
Function Calling✓ Supported

Description

Gemini 2.5 Flash is Google's price-performance, natively multimodal hybrid reasoning model. Its preview launched on April 17, 2025, followed by the stable gemini-2.5-flash endpoint on June 17, 2025. It targets low-latency, high-throughput, and large-scale agentic tasks that require reasoning, accepts text, images, video, and audio, and produces text. Its knowledge cutoff is January 2025. Google has not disclosed its parameter count, and no shutdown date for the stable endpoint had been announced as of August 2026.

Its defining innovation is hybrid reasoning: developers can turn thinking on or off and use a thinking budget to balance quality, cost, and latency. Google states that even with thinking disabled, the model maintains Gemini 2.0 Flash speed while improving performance. A million-token context plus search grounding, code execution, and function calling lets it combine large-scale content processing with tool-driven work. Shut-down preview endpoints such as gemini-2.5-flash-preview-09-2025 should not be confused with the current stable endpoint.

Key Capabilities

  • Hybrid Reasoning: Thinking can be enabled or disabled and assigned a budget to balance fast responses against deeper analysis.
  • Native Multimodal Understanding: Jointly processes text, images, video, and audio for cross-media summarization, question answering, classification, and extraction.
  • Million-Token Context: A 1,048,576-token input window accommodates large document sets, long recordings, video, and multi-file corpora.
  • Agentic Tool Use: Function calling, code execution, and Structured Outputs support calculations, business API calls, and structured workflows.
  • Grounded Retrieval: Google Search, Google Maps, File Search grounding, and URL Context incorporate current web information, places, and private files.
  • High-Throughput Content Processing: Low-latency, large-scale optimization suits classification, translation, moderation, summarization, and data transformation.
  • Multi-Level Reasoning Applications: One model can reduce thinking overhead for simple requests and allocate more reasoning to harder ones, simplifying model routing.

Technical Strengths

FeatureBenefit
Adjustable Thinking BudgetDevelopers control reasoning depth per request for finer optimization of quality, latency, and cost.
Native Multimodal ArchitectureA unified model reduces the need for separate image, audio, and video recognition pipelines in cross-media applications.
Very Large Input CapacityThe million-token window reduces fragmentation of large corpora and supports synthesis across sections and files.
Broad Grounding SupportSearch, Maps, file, and URL grounding reduce stale information and factual errors from relying only on training knowledge.
Tools and Structured OutputCode execution, function calling, and schema-constrained results connect reasoning to verifiable business actions.
Multiple Consumption OptionsBatch, Flex, and Priority inference respectively support offline throughput, elastic capacity, and prioritized latency.

Pricing

Token TypeLinkAI PriceOfficial Price
Input$0.225 / 1M tokens$0.3 / 1M tokens
Cached input$0.0225 / 1M tokens$0.03 / 1M tokens
Output$1.875 / 1M tokens$2.5 / 1M tokens
reasoning_tokens$1.875 / 1M tokens$2.5 / 1M tokens

More from Google