Volcengine

Seedance 2.0

Seedance 2.0 is a multimodal controllable video generation model developed by ByteDance’s Seed Team. Launched in early February 2026, it is now integrated into Doubao, Jimeng AI, and Volcano Engine (Model ID: doubao-seedance-2-0-260128), with an accelerated version—Seedance 2.0 Fast—available for low-latency scenarios.

Modalities
Text to Video · Image to Video
Starting price
From $3.87 / call
Calculator

Bytedance

README

Seedance 2.0 is ByteDance Seed's next-generation video creation model, released on February 12, 2026, after Seedance 1.0 and Seedance 1.5 Pro. It uses a unified, efficient, large-scale architecture for joint multimodal audio-video generation. The model accepts text, image, video, and audio inputs and natively produces 4–15 second audio-video clips at 480p or 720p. A Seedance 2.0 Fast variant is available for lower-latency use cases; ByteDance has not disclosed the parameter count.

Compared with 1.5 Pro, Seedance 2.0 advances from synchronized audio-video generation to unified multimodal audio-video generation. Creators can provide up to nine images, three video clips, and three audio clips at once, allowing the model to reference subjects, composition, motion, camera work, effects, rhythm, and sound. Its defining improvements center on physical plausibility in complex motion, long-range consistency, and instruction control. It also adds prompt-based editing of selected clips, characters, actions, or storylines and continuous video extension, while stereo audio generation can combine dialogue, ambience, effects, and music.

Key Capabilities

  • Unified Multimodal Reference: Mixes natural-language instructions with up to nine images, three videos, and three audio clips, drawing on subjects, composition, style, motion, camera work, effects, and sound.
  • Text/Image-to-Audio-Video: Produces complete 4–15 second clips at 480p or 720p from text or imagery, supporting both single-shot and multi-shot presentation.
  • Complex Motion and Interaction: Improves temporal stability and physical plausibility for multi-person interaction, competitive sports, and fine-grained actions, reducing structural failures and unnatural movement.
  • Director-Level Control: Plans shot scale, performance, lighting, shadows, camera movement, and narrative pacing from prompts, including complex scripts and composed camera sequences.
  • Instruction-Based Video Editing: Modifies selected clips, characters, actions, or plot elements while aiming to preserve unaffected content and subject identity.
  • Video Extension: Generates a continuous follow-on shot from an existing video and a new prompt, extending subjects, visual style, action logic, and story direction.
  • Stereo Joint Audio: Generates dialogue or narration, environmental effects, and background music in sync with visual events and editing rhythm.

Technical Strengths

FeatureBenefit
Unified Audio-Video ArchitectureModels visuals and sound within one generation system, reducing the semantic and timing mismatch associated with attaching audio in a separate pass.
Four-Modal Mixed ConditioningCombines textual intent with image, video, and audio assets so characters, style, motion, camera work, and sound can all serve as explicit references.
Sparse ArchitectureIdentified by ByteDance as a source of computational efficiency, making large-scale multimodal modeling more practical for content production.
Joint Multimodal TrainingShares knowledge and representations across tasks, improving generalization to new combinations of references and complex creative instructions.
Long-Range Consistency OptimizationStrengthens continuity of subjects, voices, action logic, and narrative across shots, improving the usability of 15-second multi-shot content.
Stereo Multi-Track ExpressionCoordinates voice, ambient effects, and music with visual timing, reducing the amount of downstream sound-design work required.

Pricing

How the estimate works

Video tokens equal (input video duration + output video duration) × output width × output height × 24 FPS ÷ 1,024. Without video input, input duration is zero.

ResolutionvideosToken TypeLinkAI PriceOfficial Price
1080PfalseOutput$6.93 / 1M tokens$7.7 / 1M tokens
1080PtrueOutput$4.23 / 1M tokens$4.7 / 1M tokens
480PfalseOutput$6.3 / 1M tokens$7 / 1M tokens
480PtrueOutput$3.87 / 1M tokens$4.3 / 1M tokens
720PfalseOutput$6.3 / 1M tokens$7 / 1M tokens
720PtrueOutput$3.87 / 1M tokens$4.3 / 1M tokens

More from Bytedance