Gemini

Gemini 3.8 Flash

Google's high-performance multimodal model for long-horizon software engineering, autonomous agents, and complex enterprise workflows, with million-token context and broad tooling.

Modalities
Chat
Starting price
From $3.75 / 1M out
Context
1.0M context

Google

README

Supported Functionality

ItemSpecification
InputText, image, video, audio, PDF
OutputText
Context1,048,576 tokens
Max Output65,536 tokens
Vision✓ Supported
Function Calling✓ Supported

Description

Gemini 3.8 Flash is a generally available, stable multimodal model updated by Google in September 2026 under the model code gemini-3.8-flash. It accepts text, image, video, audio, and PDF inputs, produces text, and provides a 1,048,576-token input window with up to 65,536 output tokens.

The model targets long-horizon software engineering, autonomous agents, and complex enterprise workflows, using smaller reasoning steps to call tools and verify intermediate results throughout a task. It supports low, medium, and high thinking levels with medium as the default, along with caching, code execution, File Search, function calling, Google Search and Maps grounding, structured output, and URL Context. Computer Use is in Preview; image generation, audio generation, and the Live API are not supported.

Key Capabilities

  • Long-Horizon Software Engineering: Diagnoses issues, performs multi-file refactoring, runs tests, and iterates on large real-world repositories while maintaining task continuity.
  • Autonomous Agents: Plans multi-step work, orchestrates functions and built-in tools, handles environmental feedback, and verifies intermediate outcomes.
  • Native Multimodal Understanding: Analyzes text, images, video, audio, and PDFs together to synthesize evidence across media.
  • Code Execution and Tool Use: Connects reasoning with code execution, file retrieval, custom functions, and external systems in verifiable loops.
  • Complex Enterprise Workflows: Handles analysis, extraction, decision support, and automation across documents, systems, and large data pipelines.
  • Grounding and External Information: Combines Google Search, Maps grounding, and URL Context for answers that require current web or location evidence.
  • Structured and Adjustable Reasoning: Constrains responses with structured output and selects an appropriate thinking level for each task.

Technical Strengths

FeatureBenefit
1,048,576-Token ContextKeeps large repositories, long videos, multiple PDFs, and cross-system material within one task with less fragmentation.
65,536-Token Maximum OutputSupports long analyses, multi-file change plans, and structured results spanning several execution stages.
Three Thinking Levelslow, medium, and high align reasoning depth with real-time interaction, general agents, and deep analysis; medium is the default.
Broad Built-In Tool SupportCode execution, File Search, Search, Maps, URL Context, and function calling can be composed into complete workflows.
Five Native Input TypesA single model understands text, images, video, audio, and PDFs for cross-media agents and document systems.
Stable Production VersionA fixed GA model code supports version locking, regression testing, continuous deployment, and Google's managed-agent workflows.

Frequently Asked Questions

When should I use low, medium, or high thinking with Gemini 3.8 Flash?

Start with low for real-time answers, incident response, drafts, and fast analysis; use the default medium for complex coding and general agent work; evaluate high for mathematics, deep analysis, and difficult multi-step tasks. Compare first-pass completion, tool-call count, latency, and failure recovery on representative workloads.

Is Gemini 3.8 Flash suitable for long-horizon software engineering?

Yes. Google identifies long-horizon software engineering as a core focus, including complex multi-file refactoring, real-repository issue resolution, and deterministic tool execution; applications should still include tests, static checks, and code review in the agent's validation loop.

Which multimodal inputs and outputs does Gemini 3.8 Flash support?

It accepts text, images, video, audio, and PDFs, while its output type is text. It does not directly generate images or audio and does not support the Live API, so real-time voice and media-generation workloads require the appropriate specialized models.

Gemini 3.8 Flash is the current stable model, and Google has replaced Gemini 3.7 Flash with it while automatically routing 3.7 requests to 3.8. Existing applications should explicitly change the model code to gemini-3.8-flash and rerun function-calling and structured-output regression tests to avoid relying on future routing behavior.

Which request settings should change when migrating to Gemini 3.8 Flash?

Set the model code to gemini-3.8-flash, use thinking_level with low, medium, or high, and do not send the unsupported minimal value. Remove deprecated temperature, top_p, top_k, and candidate_count, keep the final user message non-empty, and retest function calling and multimodal payloads.

How do I call Gemini 3.8 Flash on LinkModel?

Gemini 3.8 Flash is available on LinkModel. Use the public model ID, request format, and parameters shown on LinkModel's live model page, then begin with a small text request to confirm authentication, response handling, and streaming before adding multimodal input and tools.

How should I validate the Gemini 3.8 Flash integration on LinkModel?

Confirm the public ID, context limits, multimodal inputs, thinking controls, and tool support on LinkModel's current model page, then test short text, long context, image or PDF input, structured output, and function calling separately. Computer Use remains a Google Preview capability, and Google-specific fields should not be assumed to pass through LinkModel unchanged; production validation should also cover timeouts, refusals, tool failures, and retries.

Pricing

Token-based pricing

Our pricing is based on image and text token usage. The final cost depends on the tokens consumed.

Token TypeLinkAI PriceOfficial Price
Input$0.75 / 1M tokens$0.75 / 1M tokens
Cached input$0.075 / 1M tokens$0.075 / 1M tokens
Output$3.75 / 1M tokens$3.75 / 1M tokens
Reasoning output$3.75 / 1M tokens$3.75 / 1M tokens

More from Google