Volcengine

Seedance 2.0 Fast

ByteDance Seed’s low-latency joint audio-video model supports four-modal references, synchronized 4–15 second clips, and multi-shot control for rapid iteration, batch video creation, and creative previews.

Modalities
Text to Video · Image to Video
Starting price
From $2.97 / call

Bytedance

README

Seedance 2.0 Fast is ByteDance Seed’s accelerated Seedance 2.0 variant for low-latency video-generation scenarios. The Seedance 2.0 family was officially released on February 12, 2026 and uses a unified, efficient, large-scale architecture for joint multimodal audio-video generation. ByteDance’s paper defines Fast as an accelerated variant intended to improve generation speed, but the reviewed sources do not disclose a separate launch date, parameter count, or fixed speed multiplier.

The variant retains the family’s combined text, image, video, and audio reference system, with the open platform accepting up to nine images, three video clips, and three audio clips in one request. Volcano Engine documents 4–15 second output, 480p or 720p resolution, 24 fps, synchronized audio, multiple aspect ratios, and an asynchronous task workflow, making the model suitable for frequent iteration, rapid previews, and batch short-form creation.

Key Capabilities

  • Low-Latency Generation: Serves creative iteration and latency-sensitive applications that need to review a result and revise the request more quickly than a standard-generation workflow.
  • Four-Modal Referencing: Combines natural-language instructions with image, video, and audio assets to reference subjects, style, composition, action, camera movement, rhythm, and sound in one task.
  • Native Synchronized Audio-Video: Generates speech, ambience, sound effects, and background music alongside the picture, with an option to disable generated audio when a separate sound workflow is preferred.
  • Short Multi-Shot Narrative: Produces 4–15 second single-shot or multi-shot clips and responds to prompt instructions for shot scale, performance, lighting, camera movement, and narrative pacing.
  • Motion and Instruction Control: The family improves temporal stability, physical plausibility, and instruction response for multi-person interaction, sports, fine-grained action, and combined camera movements.
  • Continuous Clip Workflow: The API can return the final frame of a generated video so it can be used as the first frame of a follow-on task, supporting the rapid construction of connected clips.

Technical Strengths

FeatureBenefit
Unified Multimodal Audio-Video ArchitectureModels visuals and sound in one system, helping actions, dialogue, effects, and music remain semantically and temporally aligned.
Sparse-Architecture EfficiencyByteDance identifies sparse architecture as one source of the family’s computational efficiency, supporting large-scale multimodal generation and faster iteration.
Joint Multimodal TrainingShares representations and knowledge across text, image, video, and audio, improving generalization to new asset combinations and compound instructions.
Configurable Standardized OutputUses integer durations, 480p/720p output, and several landscape, portrait, and square ratios for common short-form delivery formats, including an adaptive ratio option.
Optional Synchronized AudioLets developers produce a complete audiovisual preview or generate silent visuals when dialogue and sound will be handled in a separate post-production pipeline.
Asynchronous Tasks and Reproducible SeedsSeparates task submission from result retrieval for batch management and supports a random seed for recording experiment conditions and controlled iteration.

Frequently Asked Questions

How much faster is the Fast variant than the standard model?

ByteDance describes it as an accelerated variant for low-latency use cases but does not publish one fixed multiplier that applies across every duration, resolution, and input type. Actual latency also depends on task length, reference assets, platform queues, and service configuration, so teams should benchmark their own representative requests.

How should multimodal references be organized?

Give every asset one explicit job, such as character appearance, composition, camera movement, action rhythm, or sound. Refer to each asset in the prompt and state both what should be copied and what should not transfer, reducing conflicts between identities, styles, and movements.

How can dialogue and sound effects be made more controllable?

Place spoken lines in quotation marks and describe the speaker, timing, environmental sound, and background music separately. If the workflow already has dedicated dubbing and mixing stages, disable generated audio to avoid overlapping it with the post-production soundtrack.

What are the documented Seedance 2.0 Fast API limitations?

Seedance 2.0 Fast currently supports 480p and 720p but not 1080p, camera_fixed, or the offline inference service tier. Workflows that require higher-resolution delivery, a parameterized fixed-camera flag, or offline batch processing should verify another Seedance version’s current API before migration.

Can the Volcano Engine model ID be used directly on LinkModel?

Not by default. doubao-seedance-2-0-fast-260128 is a Volcano Engine-specific identifier, while LinkModel defines its own model IDs; do not place the provider identifier in a LinkModel request unless LinkModel explicitly documents that exact value.

Pricing

videosToken TypeLinkAI PriceOfficial Price
falseoutput$5.040000 / 1M tokens$5.600000 / 1M tokens
trueoutput$2.970000 / 1M tokens$3.300000 / 1M tokens

More from Bytedance