Gemini Nano Banana 2.1
Google's efficient image generation and conversational editing model with 1K–4K output, multi-image fusion, character consistency, accurate text, and search grounding.
- Modalities
- Text to Image · Image to Image
- Starting price
- From $1.125 / 1M in
README
Gemini Nano Banana 2.1 is Google's efficient image generation and conversational editing model, updated in October 2026 with the stable model code gemini-nano-banana-2.1. It accepts text, images, video, and PDFs and returns images and text, with a 131,072-token input limit and a 32,768-token output limit. It is the updated version of Nano Banana 2, also known as Gemini 3.1 Flash Image.
The model improves visual quality at 1K, 2K, and 4K, prompt adherence, multi-turn character consistency, text rendering, and infographic layout while retaining Flash-level efficiency. It can combine up to 14 reference images, preserve up to four characters and ten objects, use Google Web and Image Search grounding, and operate at minimal, medium, or high thinking levels.
Key Capabilities
- Text-to-Image Generation: Creates photography, illustration, posters, infographics, and design assets from natural-language prompts.
- Conversational Image Editing: Adds, removes, or changes elements over multiple turns while preserving established composition.
- Multi-Image Fusion: Uses up to 14 references in one workflow to combine people, products, styles, and scenes.
- Character and Object Consistency: Preserves the identity of up to four characters and visual fidelity for up to ten objects.
- Multi-Resolution Output: Generates at 1K, 2K, or 4K with 1K as the default and improved wide and panoramic rendering.
- Text and Infographic Rendering: Improves the accuracy and legibility of titles, labels, copy, and infographic layouts.
- Search-Grounded Generation: Uses Google Web and Image Search for current textual and visual evidence before rendering.
Technical Strengths
| Feature | Benefit |
|---|---|
| 131K Multimodal Input | Uses text, images, video, and PDFs as generation context for richer creative workflows. |
| 14-Image Reference Fusion | Combines complex character, product, and visual-style references with less manual compositing. |
| Explicit Consistency Limits | Supports up to four characters and ten objects, making planning clearer for series and product visuals. |
| 1K–4K Output | Covers quick previews and high-resolution delivery while improving wide and panoramic aspect ratios. |
| Three Thinking Levels | Uses minimal, medium, or high to adjust visual reasoning for different compositions. |
| Search Grounding and SynthID | Grounds factual visuals in search and marks every generated image with SynthID for provenance. |
Frequently Asked Questions
How does Gemini Nano Banana 2.1 differ from Nano Banana 2?
Nano Banana 2.1 is the updated version of Nano Banana 2 (gemini-3.1-flash-image) and Google's recommended efficient workhorse for new projects. It improves visual quality, prompt adherence, text rendering, multi-turn consistency, and search grounding, and uses the distinct model code gemini-nano-banana-2.1.
How should I maintain character consistency?
The model supports consistency for up to four characters, making it useful for storyboards, campaign series, and small casts. Reuse clear references and restate names, appearance, wardrobe, framing, and locked traits on each turn; exceeding the documented count increases the risk of identity mixing.
How can I improve text in posters and infographics?
Finalize the copy first, then provide the exact wording, language, hierarchy, position, and layout requirements. For complex infographics, establish the structure first and correct titles, labels, and data through follow-up edits. Manually proofread spelling, numbers, and brand details before publication.
When should I use minimal, medium, or high thinking?
Start with minimal for straightforward generation and local edits, and use the default medium for multi-object scenes, infographics, or reference fusion. Test high for complex spatial relationships, grounded content, or stricter composition, comparing adherence, detail, and processing time on the same source set.
How do I call Gemini Nano Banana 2.1 on LinkModel?
Gemini Nano Banana 2.1 is available on LinkModel. Open its live model page, copy the published model ID, and submit text, reference images, or other supported input with the request structure currently shown there. Use the live page for platform fields, image response format, resolution, and thinking levels.
How do I validate Gemini Nano Banana 2.1 on LinkModel?
Complete a basic text-to-image request with the ID shown on LinkModel, then separately test editing, multi-turn continuity, 14 references, 1K/2K/4K output, text rendering, and search grounding. Record response format, image count, resolution, character drift, and errors, and confirm which thinking levels, video/PDF inputs, and Batch features LinkModel currently exposes.
Pricing
Token-based pricing
Our pricing is based on image and text token usage. The final cost depends on the tokens consumed.
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| Input | $1.125 / 1M tokens | $1.5 / 1M tokens |
| Cached input | $0.1125 / 1M tokens | $0.15 / 1M tokens |
| Image · Output | $22.5 / 1M tokens | $30 / 1M tokens |
| Text · Output | $5.625 / 1M tokens | $7.5 / 1M tokens |
| Reasoning output | $5.625 / 1M tokens | $7.5 / 1M tokens |