Gemini

Gemini 3.1 Flash Lite

An ultra-low-cost, ultra-fast native four-modal large model developed by Google DeepMind, equipped with adaptive dynamic reasoning, million-token long context and high-throughput batch tool calling, optimized for massive high-frequency industrial workloads including real-time customer service, bulk text classification & translation, mass document summarization, lightweight agents, receipt data extraction and high-concurrency consumer chat applications.

Modalities
Chat
Starting price
From $1.125 / 1M out
Context
100K context

Google

README

Supported Functionality

ItemSpecification
InputText, image, video, audio, PDF
OutputText
Context1,048,576 tokens
Max Output65,536 tokens
Vision✓ Supported
Function Calling✓ Supported

Description

Gemini 3.1 Flash-Lite is Google's lightweight, natively multimodal reasoning model in the Gemini 3 family. Its preview launched on March 3, 2026, followed by the stable gemini-3.1-flash-lite endpoint on May 7, 2026. Based on Gemini 3 Pro, it targets high-throughput, latency-sensitive, and cost-sensitive tasks, accepts text, image, video, audio, and PDF inputs, and has a January 2025 knowledge cutoff. Google has not disclosed its parameter count.

Its main advance is bringing adjustable reasoning and native multimodal processing to a lightweight price tier. Google reports high-thinking scores of 86.9% on GPQA Diamond, 76.8% on MMMU-Pro, and 72.0% on LiveCodeBench. Citing Artificial Analysis, Google also reports 2.5 times faster time to first answer token and 45% higher output speed than Gemini 2.5 Flash. The stable endpoint has an earliest shutdown date of May 7, 2027, with gemini-3.5-flash-lite listed as the recommended replacement.

Key Capabilities

  • Efficient Reasoning: Supports minimal, low, medium, and high thinking levels, allowing developers to balance latency for bulk tasks against deeper reasoning for harder requests.
  • Native Multimodal Understanding: Jointly analyzes text, images, video, audio, and PDFs, scoring 76.8% on MMMU-Pro for cross-media classification, labeling, and extraction.
  • Million-Token Context: Accepts up to 1,048,576 tokens for large document sets, long recordings, video, and cross-file synthesis, although detailed recall at extreme lengths should be evaluated per workload.
  • Lightweight Agents: Function calling, code execution, structured outputs, and URL Context support query routing, data processing, and multi-step business automation.
  • Grounded Retrieval: Google Search, Google Maps, and File Search grounding enable responses informed by current web information, location data, and private corpora.
  • Multilingual Processing: An 88.9% MMMLU score supports high-volume translation, moderation, customer-message classification, and cross-language question answering.
  • Fast Code Generation: A 72.0% LiveCodeBench score supports UI, dashboard, and routine script generation, while complex software engineering is better assigned to a higher-tier model.

Technical Strengths

FeatureBenefit
Low-Latency InferenceGoogle's measured output speed of 363 tokens per second helps reduce interactive wait times and raise batch throughput.
Adjustable Thinking LevelsDevelopers can match reasoning depth to task difficulty for finer control over quality, latency, and cost.
Native Multimodal ArchitectureA unified model reduces the need for separate image, audio, and video recognition pipelines in cross-media applications.
Very Large Input CapacityThe million-token window reduces fragmentation of large corpora and supports synthesis across sections and files.
Broad Tool EcosystemFunction calling, grounded search, file retrieval, and code execution connect language understanding to real business operations.
Scalable Inference OptionsBatch, Flex, and Priority inference let teams optimize for offline throughput, elastic capacity, or prioritized latency.

Pricing

Token-based pricing

Our pricing is based on image and text token usage. The final cost depends on the tokens consumed.

Token TypeLinkAI PriceOfficial Price
Input$0.1875 / 1M tokens$0.25 / 1M tokens
Cached input$0.01875 / 1M tokens$0.025 / 1M tokens
Output$1.125 / 1M tokens$1.5 / 1M tokens
Reasoning output$1.125 / 1M tokens$1.5 / 1M tokens

More from Google