MODEL LIBRARY · ONE API

Every frontier model. One integration.

Compare leading image, video, and language models— then ship with one API key.

57 models
Claude Opus 5.5
AnthropicChat

Anthropic's high-capability reasoning model for long-running coding agents and professional knowledge work, with million-token context, vision, and sustained tool use.

Starts from$17 / 1M outTry it
GPT 6 Luna
OpenAIChat

OpenAI's efficient reasoning model for focused high-volume work, extraction, and frequent automation, with million-token context, vision, and tool support.

Starts from$0.375 / 1M outTry it
GPT 6 Sol
OpenAIChat

OpenAI's general-purpose reasoning model for complex coding, agent workflows, and professional work, with million-token context, vision, and broad tool support.

Starts from$7.5 / 1M outTry it
GPT Image 2.5 Sunburst
OpenAIText-to-Image

A precision-focused visual generation and editing model supporting multiple references, mask-guided changes, custom dimensions, and transparent backgrounds for demanding creative production.

Starts from$0.00441 / imageTry it
GPT Image 2.5 Flare
OpenAIText-to-Image

A fast, high-quality image generation and editing model with reference-based changes, custom dimensions, multiple quality settings, and transparent output for frequent visual production.

Starts from$0.00441 / imageTry it
GPT 6 Astra
OpenAIChat

OpenAI's flagship reasoning model for demanding end-to-end work. It supports an approximately 1.05-million-token context, up to 128K output tokens, and five reasoning-effort levels for complex analysis, coding, planning, and professional text generation.

Starts from$37.5 / 1M outTry it
Claude Fable 5.1
AnthropicChat

Anthropic's frontier language model for long-running coding agents, multistep research, and complex document analysis in demanding knowledge workflows.

Starts from$42.5 / 1M outTry it
MiniMax H3
MiniMaxText-to-Video

MiniMax's omni-modal audio-video generation system unifies text, image, video, and audio context to create videos up to 15 seconds at 2K with native stereo sound for advertising, e-commerce, and film production.

Starts from$0.08 / secondTry it
Nano Banana 2 Lite
GoogleText-to-Image

Google's efficiency-tier image model for ultra-low latency and low cost, supporting text-to-image, interleaved generation-editing, and multi-turn local edits with SynthID+C2PA watermarks; ideal for high-volume interactive apps and prototyping.

Starts from$0.0252 / imageTry it
Seedance 2.5
BytedanceText-to-Video

ByteDance’s production-oriented multimodal video model for generating and editing coherent clips up to 30 seconds, with synchronized audio, timestamp-level control, and up to 30 image, 10 video, and 10 audio references for stronger character, scene, motion, and style consistency.

Starts from$0.09766 / secondTry it
Claude Fable 5
AnthropicChat

Anthropic's frontier language model with a one-million-token context, long-running agentic coding, and multistep reasoning for complex research and knowledge work.

Starts from$42.5 / 1M outTry it
Nano Banana 2
GoogleText-to-Image

Nano Banana (gemini-3.1-flash-image) is a highly multifunctional generative model designed for speed, accuracy, and creative flexibility. It excels at transforming text prompts into stunning visual effects while providing advanced composition and style transfer capabilities.

Starts from$0.03375 / imageTry it
HappyHorse 1.0
AlibabaText-to-Video

Alibaba ATH’s 15B video generation model with native audio-video generation, supporting dialogue, ambient sound, and lip sync. Ranked No. 1 on both Artificial Analysis Video Arena leaderboards, it is ideal for short dramas, ads, and dialogue-driven content.

Starts from$0.1158 / secondTry it
Kling V3
KlingText-to-Video

Kuaishou's unified multimodal video model with native 4K/60fps output, AI Director multi-shot storyboarding, multilingual native audio, and ultimate character consistency, unifying video understanding, generation, and editing in one workflow.

Starts from$0.06097 / secondTry it
Gemini 3.1 Pro Preview
GoogleChat

Google's frontier reasoning model, leading on novel logic (ARC-AGI-2 77.1%), graduate-level science (GPQA Diamond 94.3%) and agentic reliability, with a 1M-token context and full multimodal input — built for complex software engineering, long-document analysis and autonomous agent workflows.

Starts from$9 / 1M outTry it
Seedream 5.0 Lite
BytedanceText-to-Image

ByteDance's intelligent image model with Chain of Thought visual reasoning and real-time web search, comprehensively upgraded in understanding, reasoning, and generation, supporting up to 4K output and 14-image reference editing.

Starts from$0.030305 / imageTry it
DeepSeek V4 Flash
DeepSeekChat

DeepSeek's open-source efficiency-tier flagship — 284B MoE with 13B active parameters, 1M context, dual thinking/non-thinking modes, priced at a fraction of closed frontier models for top price-performance.

Starts from$0.224 / 1M outTry it
GLM 5.3
Z.aiChat

Z.ai’s latest flagship text model for complex software engineering and long-horizon agents supports up to a 1M-token context, 128K-token output, three reasoning-effort levels, tool calling, structured output, and context caching, with stronger performance and token efficiency across coding, terminal, and multi-step professional tasks.

Starts from$3.96 / 1M outTry it
DeepSeek V4 Pro
DeepSeekChat

DeepSeek's open-weight flagship 1.6T MoE model with hybrid sparse attention, 1M-token context, competition-grade coding (LiveCodeBench 93.5), and ultra-low pricing — built for code, reasoning, and long-horizon agents.

Starts from$2.784 / 1M outTry it
Claude Opus 5
AnthropicChat

Anthropic's high-capability language model for complex coding, long-running agents, and deep reasoning with a one-million-token context for enterprise workflows.

Starts from$21.25 / 1M outTry it
Kimi K3
MoonshotChat

Moonshot AI's open flagship LLM: a 2.8T-parameter MoE (first open 3T-class model) with Kimi Delta Attention and Attention Residuals, native vision, 1M-token context, and always-on thinking for long-horizon coding, knowledge work, and reasoning.

Starts from$15 / 1M outTry it
GPT 5.6 Sol
OpenAIChat

OpenAI's flagship frontier language model designed for complex reasoning, professional knowledge work, advanced agents, and software engineering, featuring long-context processing, multimodal understanding, and powerful tool capabilities.

Starts from$15 / 1M outTry it
GPT 5.6 Luna
OpenAIChat

OpenAI's cost-optimized general-purpose language model designed for high-volume AI workloads, supporting long-context reasoning, multimodal understanding, and agent tooling for enterprise automation, assistants, content generation, and scalable applications.

Starts from$0.9 / 1M outTry it
GPT 5.6 Terra
OpenAIChat

OpenAI's next-generation general-purpose model balancing intelligence and cost, featuring long-context reasoning, multimodal understanding, and full agent tooling for enterprise AI, coding, automation, and professional knowledge work.

Starts from$9 / 1M outTry it
Seedream 5.0 Pro
BytedanceText-to-Image

A professional image generation and editing model from ByteDance, built for strong prompt following, text rendering, multi-image reference, and high-fidelity visual detail across commercial design and creative production.

Starts from$0.045 / imageTry it
Claude Sonnet 5
AnthropicChat

Anthropic's balanced language model for agentic coding, tool use, and computer operation, with a one-million-token context for everyday development and enterprise automation.

Starts from$8.5 / 1M outTry it
GLM 5.2
Z.aiChat

Z.ai’s flagship text model for long-horizon work supports up to a 1M-token context and 128K-token output, with adjustable reasoning, tool calling, structured output, and context caching for project-scale code understanding, complex refactoring, and persistent agent workflows.

Starts from$3.96 / 1M outTry it
Seedance 2.0 Mini
BytedanceText-to-Video

ByteDance’s lightweight joint audio-video model supports four-modal references, synchronized 4–15 second clips, and video editing for e-commerce, social media, and high-volume production.

Starts from$0.030264 / secondTry it
Nano Banana Pro
GoogleText-to-Image

Google's flagship image generation and editing model built on Gemini 3 Pro, featuring up to 4K resolution, precise multilingual text rendering, real-time Google Search grounding, and studio-quality creative controls.

Starts from$0.1005 / imageTry it
Claude Opus 4.8
AnthropicChat

Anthropic's flagship LLM for agentic coding, multidisciplinary reasoning, and knowledge work, with sharper judgment and stronger honesty, a 1M-token context window, and fast mode—built for code migration, deep research, and legal/financial analysis.

Starts from$21.25 / 1M outTry it
Gemini 3.5 Flash
GoogleChat

A frontier-level ultra-fast native five-modal large model developed by Google DeepMind, equipped with dynamic multi-level deep thinking, stable agent orchestration, flagship coding capability and million-token long context, delivering frontier intelligence at half the cost, suitable for enterprise-scale agent clusters, full-stack R&D, bulk multimedia processing, massive document governance and high-concurrency commercial API scenarios.

Starts from$6.75 / 1M outTry it
Gemini 3.1 Flash Lite
GoogleChat

An ultra-low-cost, ultra-fast native four-modal large model developed by Google DeepMind, equipped with adaptive dynamic reasoning, million-token long context and high-throughput batch tool calling, optimized for massive high-frequency industrial workloads including real-time customer service, bulk text classification & translation, mass document summarization, lightweight agents, receipt data extraction and high-concurrency consumer chat applications.

Starts from$1.125 / 1M outTry it
GPT 5.5
OpenAIChat

OpenAI's natively omnimodal, agentic model processing text, image, audio, and video end-to-end, with a 1M-token context window and native computer use, leading coding and knowledge-work benchmarks for autonomous agents and research.

Starts from$22.5 / 1M outTry it
Kimi K2.6
MoonshotChat

Moonshot AI's open-weight 1T-parameter MoE multimodal agentic model, excelling in long-horizon coding, 300-sub-agent swarm orchestration, and native tool use — built for autonomous software engineering and agent workflows.

Starts from$3.4 / 1M outTry it
GPT Image 2
OpenAIText-to-Image

OpenAI's latest image generation and editing model with a reasoning thinking mode achieving 99%+ text rendering accuracy, supporting up to 2K resolution, flexible aspect ratios, and multilingual text, deeply integrated into ChatGPT and API.

Starts from$0.00441 / imageTry it
Claude Opus 4.7
AnthropicChat

Anthropic's flagship LLM excelling at long-horizon agents and coding, scoring 87.6% on SWE-bench Verified, with 1M context, adaptive thinking, and high-res vision — for complex coding and enterprise workflows.

Starts from$18.75 / 1M outTry it
GLM 5.1
Z.aiChat

Z.ai's open-source flagship LLM with 754B MoE architecture, purpose-built for agentic coding and long-horizon tasks, capable of 8-hour autonomous execution and topping SWE-Bench Pro.

Starts from$3.96 / 1M outTry it
Seedance 2.0 Fast
BytedanceText-to-Video

ByteDance Seed’s low-latency joint audio-video model supports four-modal references, synchronized 4–15 second clips, and multi-shot control for rapid iteration, batch video creation, and creative previews.

Starts from$0.048422 / secondTry it
MiniMax M2.7
MiniMaxChat

MiniMax's self-evolving 230B MoE model (10B active) with native Agent Teams, dynamic tool search, and end-to-end engineering delivery — MLE Bench Lite 66.6% trails only Opus 4.6/GPT-5.4, built for SRE, ML research, and office automation.

Starts from$0.9 / 1M outTry it
GPT 5.4 Mini
OpenAIChat

OpenAI's most capable mini model yet for coding, computer use, and subagents; supports text and image input, a 400K context window, and extensive tool calling, optimized for low-latency, high-volume production workloads.

Starts from$3.375 / 1M outTry it
GPT 5.4 Pro
OpenAIChat

OpenAI's enhanced-reasoning tier built on GPT-5.4's architecture, applying extra compute for higher accuracy on complex, high-stakes tasks, with a million-token context and deep chain-of-thought for professional research and long documents.

Starts from$135 / 1M outTry it
GPT 5.4
OpenAIChat

OpenAI's unified flagship merging Codex and GPT, the first mainline reasoning model with frontier coding capability, featuring a million-token context and native computer use, leading knowledge-work and coding benchmarks for agentic automation.

Starts from$11.25 / 1M outTry it
Claude Sonnet 4.6
AnthropicChat

Anthropic's high-value mid-tier model approaching Opus-level performance on coding, computer use, and agent orchestration, with 1M context and adaptive thinking — built for scaled coding, desktop automation, and enterprise agents.

Starts from$11.25 / 1M outTry it
Seedance 2.0
BytedanceText-to-Video

Seedance 2.0 is a multimodal controllable video generation model developed by ByteDance’s Seed Team. Launched in early February 2026, it is now integrated into Doubao, Jimeng AI, and Volcano Engine (Model ID: doubao-seedance-2-0-260128), with an accelerated version—Seedance 2.0 Fast—available for low-latency scenarios.

Starts from$0.060527 / secondTry it
GPT 5.3 Codex
OpenAIChat

OpenAI's agentic coding model merging prior Codex engineering strength with GPT-5.2's reasoning, supporting real-time mid-task steering, setting new highs on terminal and computer-use benchmarks for long-horizon autonomous development work.

Starts from$10.5 / 1M outTry it
Claude Opus 4.6
AnthropicChat

Anthropic's frontier flagship model built for deep reasoning, long-horizon agentic coding, and enterprise knowledge work — tops Terminal-Bench 2.0 and HLE with 1M context and adaptive thinking for the hardest tasks.

Starts from$18.75 / 1M outTry it
Gemini 3 Flash Preview
GoogleChat

A cost-efficient ultra-fast multimodal large model developed by Google DeepMind, equipped with dynamically adjustable deep reasoning, native four-modal parsing, million-token long context and stable batch tool calling, optimized for low-latency high-concurrency scenarios including enterprise agents, high-frequency API services, coding assistance, bulk audio-video processing and consumer chat applications。

Starts from$2.25 / 1M outTry it
Wan 2.6
AlibabaText-to-Video

Alibaba's multimodal video generation model series supporting role-play (reference-to-video), multi-shot narrative, audio-visual sync, and up to 15-second output, enabling creators to star in AI videos with their own appearance and voice.

Starts from$0.06097 / secondTry it
GPT 5.2
OpenAIChat

OpenAI's flagship reasoning model for professional work and long-running agents, setting new highs in knowledge work, coding, and long-context understanding, offered in Instant, Thinking, and Pro modes for enterprise automation.

Starts from$10.5 / 1M outTry it
Seedream 4.5
BytedanceText-to-Image

ByteDance's unified image generation and editing model with precise text rendering, native 4K resolution, and multi-image reference editing, supporting professional typography, material fidelity, and brand visual consistency for commercial design.

Starts from$0.02904 / imageTry it
MiniMax Hailuo 2.3
MiniMaxText-to-Video

MiniMax's upgraded video model built on Hailuo 02, enhancing complex body motion, facial micro-expressions, and physical realism, with expanded style support including anime, ink wash, and game CG — better quality at the same price.

Starts from$0.224 / videoTry it
Nano Banana
GoogleText-to-Image

Nano Banana (`gemini-2.5-flash-image`) is Google's natively multimodal image generation and editing model designed for speed and efficiency, supporting text-to-image, conversational editing, multi-image composition, and character consistency.

Starts from$0.02925 / imageTry it
Seedream 4.0
BytedanceText-to-Image

ByteDance's multimodal image creation engine unifying text-to-image generation and editing in one architecture, supporting up to 4K output, 30+ art style switching, and multi-reference batch generation with 10x faster inference.

Starts from$0.0232 / imageTry it
GPT 5
OpenAIChat

OpenAI's unified AI system combining a fast-response model and a deep-reasoning model via real-time routing, delivering state-of-the-art coding, math, multimodal understanding, and health reasoning for chat, development, and agentic tasks.

Starts from$7.5 / 1M outTry it
MiniMax Hailuo 02
MiniMaxText-to-Video

MiniMax's cinematic AI video generation model built on its proprietary NCR architecture with 2.5x efficiency gains, supporting native 1080p output, extreme physics simulation, and precise instruction following, ranked among the top global video models.

Starts from$0.08 / videoTry it
Gemini 2.5 Flash
GoogleChat

Google's next-gen lightweight multimodal model featuring ultra-fast inference, native multimodal fusion, and massive context capacity, optimized for high-frequency interactions, real-time responses, and cost-efficient data processing.

Starts from$1.875 / 1M outTry it
Wan3.0 Video
AlibabaText-to-Video

Alibaba’s Wan audiovisual generation model. LinkModel’s initial release is planned for text-to-video and first/last-frame generation, with up to 1080P, 30-second output and synchronized audio, subject to the capabilities enabled on the model page.

Starts from$0.05 / secondTry it