Volcengine

Bytedance/seedance-2-5

From $6.08/ call

ByteDance’s production-oriented multimodal video model for generating and editing coherent clips up to 30 seconds, with synchronized audio, timestamp-level control, and up to 30 image, 10 video, and 10 audio references for stronger character, scene, motion, and style consistency.

Text to VideoImage to Video

More from Bytedance

README

Bytedance/seedance-2-5

Seedance 2.5 is ByteDance’s multimodal video generation and editing model designed for longer, more controllable audiovisual creation. It supports text-guided generation, first- and last-frame workflows, and reference-driven creation using images, videos, and audio, producing coherent clips up to 30 seconds with synchronized sound.

Compared with Seedance 2.0, the central upgrade is not simply longer output, but a more production-oriented workflow. A single request can use up to 30 image references, 10 video references, and 10 audio references to guide character identity, product appearance, visual style, camera language, motion, and sound. Audio can also serve as the sole reference, enabling music-, dialogue-, or sound-driven video creation without requiring an image or video input.

Timestamp instructions give creators more precise control over when actions, camera changes, and scene transitions occur. Seedance 2.5 also improves complex motion, physical realism, cross-shot continuity, multi-character separation, and voice fidelity, helping longer sequences remain visually and narratively coherent. Its broader workflow supports localized video editing, extension, camera-angle changes, seamless transitions, storyboard references, and production-previsualization controls.

Key Capabilities

  • Up to 30-Second Generation: Creates longer narrative sequences with multiple actions, camera movements, pacing changes, and scene transitions in a single generation.
  • High-Capacity Multimodal Reference: Accepts up to 30 images, 10 videos, and 10 audio clips to guide characters, products, environments, motion, style, and sound.
  • Audio-Driven Creation: Supports audio-only references and synchronized audiovisual generation for dialogue, music, sound effects, and rhythm-led content.
  • Timestamp-Level Control: Lets prompts assign actions, camera changes, and transitions to specific moments within the generated video.
  • Improved Subject Consistency: Better preserves character identity, clothing, product details, visual style, and individual facial features across longer or multi-character scenes.
  • Complex Motion and Camera Control: Handles expressive subject motion, long-tail actions, camera movement, and scene progression with improved physical coherence.
  • Iterative Video Workflows: Supports editing and extension workflows that modify selected elements while preserving the surrounding composition and motion.
  • Production Previsualization: Works with storyboard and white-model references to guide composition, blocking, camera paths, and movement before final production.

Technical Strengths

FeatureBenefit
Extended temporal generationSupports more complete stories and product demonstrations without assembling many short clips
Multimodal reference capacityKeeps characters, products, environments, style, movement, and sound aligned with supplied assets
Timestamp instruction followingProvides finer control over the timing and sequence of actions, shots, and transitions
Unified audiovisual creationReduces the work required to synchronize dialogue, sound effects, music, and visual events
Improved multi-subject consistencyHelps different characters retain independent facial features, voices, and identities within the same scene
Consistency-preserving editingMakes localized revisions and content variations possible without rebuilding the entire video
Storyboard and spatial guidanceTranslates production plans into more controlled composition, blocking, and camera movement

Pricing

ResolutionvideosToken TypeLinkAI PriceOfficial Price
480Pfalseoutput$10.165000 / 1M tokens$10.700000 / 1M tokens
480Ptrueoutput$6.080000 / 1M tokens$6.400000 / 1M tokens
720Pfalseoutput$10.165000 / 1M tokens$10.700000 / 1M tokens
720Ptrueoutput$6.080000 / 1M tokens$6.400000 / 1M tokens