Gemini 2.5 Flash
Google's next-gen lightweight multimodal model featuring ultra-fast inference, native multimodal fusion, and massive context capacity, optimized for high-frequency interactions, real-time responses, and cost-efficient data processing.
- Modalities
- Chat
- Starting price
- From $0.0225 / call
- Context
- 100K context
README
Supported Functionality
| Item | Specification |
|---|---|
| Input | Text, image, video, audio |
| Output | Text |
| Context | 1,048,576 tokens |
| Max Output | 65,536 tokens |
| Vision | ✓ Supported |
| Function Calling | ✓ Supported |
Description
Gemini 2.5 Flash is Google's price-performance, natively multimodal hybrid reasoning model. Its preview launched on April 17, 2025, followed by the stable gemini-2.5-flash endpoint on June 17, 2025. It targets low-latency, high-throughput, and large-scale agentic tasks that require reasoning, accepts text, images, video, and audio, and produces text. Its knowledge cutoff is January 2025. Google has not disclosed its parameter count, and no shutdown date for the stable endpoint had been announced as of August 2026.
Its defining innovation is hybrid reasoning: developers can turn thinking on or off and use a thinking budget to balance quality, cost, and latency. Google states that even with thinking disabled, the model maintains Gemini 2.0 Flash speed while improving performance. A million-token context plus search grounding, code execution, and function calling lets it combine large-scale content processing with tool-driven work. Shut-down preview endpoints such as gemini-2.5-flash-preview-09-2025 should not be confused with the current stable endpoint.
Key Capabilities
- Hybrid Reasoning: Thinking can be enabled or disabled and assigned a budget to balance fast responses against deeper analysis.
- Native Multimodal Understanding: Jointly processes text, images, video, and audio for cross-media summarization, question answering, classification, and extraction.
- Million-Token Context: A 1,048,576-token input window accommodates large document sets, long recordings, video, and multi-file corpora.
- Agentic Tool Use: Function calling, code execution, and Structured Outputs support calculations, business API calls, and structured workflows.
- Grounded Retrieval: Google Search, Google Maps, File Search grounding, and URL Context incorporate current web information, places, and private files.
- High-Throughput Content Processing: Low-latency, large-scale optimization suits classification, translation, moderation, summarization, and data transformation.
- Multi-Level Reasoning Applications: One model can reduce thinking overhead for simple requests and allocate more reasoning to harder ones, simplifying model routing.
Technical Strengths
| Feature | Benefit |
|---|---|
| Adjustable Thinking Budget | Developers control reasoning depth per request for finer optimization of quality, latency, and cost. |
| Native Multimodal Architecture | A unified model reduces the need for separate image, audio, and video recognition pipelines in cross-media applications. |
| Very Large Input Capacity | The million-token window reduces fragmentation of large corpora and supports synthesis across sections and files. |
| Broad Grounding Support | Search, Maps, file, and URL grounding reduce stale information and factual errors from relying only on training knowledge. |
| Tools and Structured Output | Code execution, function calling, and schema-constrained results connect reasoning to verifiable business actions. |
| Multiple Consumption Options | Batch, Flex, and Priority inference respectively support offline throughput, elastic capacity, and prioritized latency. |
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| Input | $0.225 / 1M tokens | $0.3 / 1M tokens |
| Cached input | $0.0225 / 1M tokens | $0.03 / 1M tokens |
| Output | $1.875 / 1M tokens | $2.5 / 1M tokens |
| reasoning_tokens | $1.875 / 1M tokens | $2.5 / 1M tokens |