Google/gemini-3.1-pro-preview
Google's frontier reasoning model, leading on novel logic (ARC-AGI-2 77.1%), graduate-level science (GPQA Diamond 94.3%) and agentic reliability, with a 1M-token context and full multimodal input — built for complex software engineering, long-document analysis and autonomous agent workflows.
More from Google
README
Google/gemini-3.1-pro-preview
Supported Functionality ItemSpecificationInputText, Image, Audio, Video, File, CodeOutputTextContext1,048,576 tokens (~1M)Max Output65,536 tokensVision✓ SupportedFunction Calling✓ Supported (incl. customtools agentic endpoint) Description Gemini 3.1 Pro Preview was released by Google (DeepMind) on February 19, 2026 as the core flagship reasoning model of the Gemini 3 series, positioned between the fast Gemini 3 Flash and the research-focused Gemini 3 Deep Think. On the Artificial Analysis Intelligence Index it scores 57, tied for the top with GPT-5.4, and leads 6 of the 10 evaluations feeding that index. Smart ChunksSmart Chunks As a 3.1 increment, it builds on Gemini 3 Pro with roughly a 15% quality improvement while running faster and more efficiently with fewer output tokens. The headline leap is abstract reasoning: its ARC-AGI-2 score rose from 31.1% to 77.1%, more than doubling reasoning performance, and it adds low/medium/high/max thinking levels to trade cost against quality. Google CloudGoogle Key Capabilities
Novel reasoning: ARC-AGI-2 of 77.1%, more than double Gemini 3 Pro, excelling at unseen logic patterns. Google Graduate-level science: GPQA Diamond 94.3% (no tools), currently leading all publicly-scored frontier models. Smart Chunks Autonomous software engineering: SWE-Bench Verified 80.6%, top tier. Label Rer Competitive coding: LiveCodeBench Pro at 2887 Elo. SmartScope Agentic & tool coordination: BrowseComp 85.9%, MCP Atlas 69.2%, τ²-bench Telecom 99.3%. SmartScope Long-context retrieval: MRCR v2 84.9% at 128k context. Smart Chunks Multimodal understanding: MMMU-Pro 80.5%, MMMLU 92.6%, with native image/audio/video processing. SmartScope
Technical Strengths FeatureBenefit1M-token contextLoad an entire codebase or 20+ papers in one session without chunking, cutting engineering overheadFour thinking levelsTune speed/cost vs. quality per request, saving money from simple autocomplete to complex debuggingcustomtools endpointOptimized for agentic workflows using custom tools and bash, improving multi-step reliabilityCost-efficient pricing$2/$12 per 1M tokens, cheaper than Claude Opus 4.6 and GPT-5.4 on input Smart ChunksFull multimodal inputUnified reasoning over text/image/audio/video — one model spanning many enterprise workflowsToken efficiencyFewer output tokens while delivering more reliable results, Google Cloud lowering real call cost Capability Ratings DimensionRatingNotesReasoningTop-tierLeads frontier peers on both ARC-AGI-2 and GPQA DiamondCodingExcellentTop-tier SWE-Bench and standout competitive-coding EloCreative WritingStrongStrong reasoning and instruction-following, though not writing-specializedMultimodalTop-tierNative text/image/audio/video; leads multimodal benchmarksResponse SpeedModerate~110 tokens/sec is above average, but time-to-first-token of ~25.6s is on the high end Artificial AnalysisContext WindowHuge1M tokens, larger than ~97% of comparable models Design for Online Use Cases
Autonomous coding agents: Multi-step engineering where the model reads a bug report and fixes code on its own. Scientific research synthesis: 1M context plus high GPQA enables reasoning across dozens of papers at once. Finance & spreadsheet workflows: Stronger autonomous task execution in structured domains like finance and spreadsheets. OpenRouter Long-document analysis: Process lengthy contracts, whole codebases, or large knowledge bases in a single pass. Multimodal applications: Analyzes a YouTube video with no upload required, suited to video/audio understanding pipelines. Label Rer Complex decision support: Bring disparate data into a single view, visualize complex topics, and plan.
A few notes on the data: this is a preview model with no confirmed GA date, and pricing tiers up to $4 input / $18 output per 1M tokens for long context above 200K tokens. Benchmark figures come from Google's model card and Artificial Analysis; some third-party leaderboards report GPQA Diamond slightly differently (94.1% vs 94.3%). Want me to drop this into a downloadable Word or Markdown file, or trim any section? Label Rer
Pricing
Tiered by input prompt tokens (incl. cache): once over the threshold, the whole request is billed at the higher tier.
| Token Type | Short context ≤200K | Long context >200K |
|---|---|---|
| input | $1.5$2-25% | $3$4-25% |
| output | $9$12-25% | $13.5$18-25% |
| reasoning_tokens | $9$12-25% | $13.5$18-25% |
| cache_read | $0.15$0.2-25% | $0.3$0.4-25% |