Google/gemini-2.5-flash
Google's next-gen lightweight multimodal model featuring ultra-fast inference, native multimodal fusion, and massive context capacity, optimized for high-frequency interactions, real-time responses, and cost-efficient data processing.
More from Google
README
Google/gemini-2.5-flash
Supported Functionality
Item Specification Input Text, Image, Audio, Video Output Text Context 1,000,000 tokens Max Output 8,192 tokens Vision ✓ Supported Function Calling ✓ Supported Description Developed by Google, this next-generation lightweight multimodal flagship is positioned as the "king of efficiency," balancing extreme speed with robust performance. Built on the cutting-edge Gemini 2.5 architecture, it achieves benchmark results comparable to heavyweight models while operating with a significantly smaller footprint. Its core breakthrough lies in the ultimate optimization of both speed and cost. Compared to its predecessor, it delivers a massive leap in Time to First Token (TTFB), significantly lowers operational costs, and achieves key advancements in native multimodal fusion, Needle-in-a-Haystack retrieval, and function-calling accuracy for agents.
Key Capabilities ● Ultra-fast Inference: Delivers industry-leading generation speed and low latency via architectural optimizations, built for real-time applications. ● Native Multimodal Understanding: Seamlessly processes and fuses text, image, audio, and video inputs for complex cross-modal reasoning. ● Million-scale Context Window: Natively supports up to 1M tokens, enabling single-prompt processing of huge document repositories or hours of video. ● Advanced Function Calling: Achieves exceptional API-calling and parameter-extraction accuracy, acting as the reliable brain for complex Agent workflows. ● Precise Instruction Following: Excels at adhering to system prompts and outputting strict structured data (e.g., nested JSON), simplifying developer integration.
Technical Strengths
Feature Benefit Native Multimodal Architecture Prevents information loss common in patched models, ensuring highly accurate cross-modal reasoning. High-Speed Inference Engine Delivers millisecond-level latency, dramatically improving the fluidity of chatbots and real-time apps. Sparse Attention Mechanism Significantly reduces memory and compute costs for ultra-long contexts, boosting concurrency under heavy loads. Dynamic Token Routing Intelligently allocates compute resources to maintain quality while offering an exceptional price-performance ratio. Enhanced Instruction Tuning Makes the model highly stable and predictable in strict development environments and multi-step agent orchestration. Enterprise-grade Security Built-in with Google's latest multi-dimensional guardrails to prevent malicious injections and ensure compliance. Capability Ratings
Dimension Rating Notes Reasoning Strong Demonstrates solid performance in complex logic and common-sense reasoning for most business needs. Coding Strong Proficient in generating and debugging multiple languages, supporting extensive codebase analysis. Creative Writing Moderate Fluent and natural, though slightly less nuanced in deep literary styles compared to ultra-large models. Multimodal Top-tier Accurately identifies video timeline events, audio nuances, and highly complex charts. Response Speed Very Fast One of the fastest native multimodal models in the industry with ultra-low latency and high throughput. Context Window Huge Natively supports 1M tokens for seamless and lossless processing of massive corpora. Use Cases ● Real-time Customer Service Agents: Leverages ultra-low latency and multimodality to build highly responsive, human-like voice or video bots. ● Massive Document Summarization: Processes hundreds of financial reports or legal files at once for accurate cross-document retrieval and comparison. ● Automated Video Analysis: Ingests long videos directly to auto-generate timeline markers, summaries, or answer detail-specific visual questions. ● High-frequency Data Processing: Executes large-scale unstructured data cleaning, classification, and JSON structuring at extremely low cost. ● Complex Workflow Automation: Acts as a central hub seamlessly connecting external APIs and internal enterprise systems via robust function calling.
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| input | $0.225000 / 1M tokens | $0.300000 / 1M tokens |
| output | $1.875000 / 1M tokens | $2.500000 / 1M tokens |
| reasoning_tokens | $1.875000 / 1M tokens | $2.500000 / 1M tokens |
| cache_read | $0.022500 / 1M tokens | $0.030000 / 1M tokens |