Volcengine

Seedance 2.0 Mini

ByteDance’s lightweight joint audio-video model supports four-modal references, synchronized 4–15 second clips, and video editing for e-commerce, social media, and high-volume production.

Modalities
Text to Video · Image to Video
Starting price
From $1.89 / call

Bytedance

README

Seedance 2.0 Mini is a lightweight Seedance 2.0 video-generation tier from ByteDance’s Doubao foundation-model team. BytePlus’s official model catalog records an update date of June 15, 2026. It follows the Seedance 2.0 family’s unified multimodal audio-video generation direction and is positioned for broader everyday creation, rapid experimentation, and frequent production workloads. The reviewed official sources do not disclose its parameter count or separate architectural details.

BytePlus lists Mini as a basic Seedance 2.x model supporting text-to-video, first-frame and first-plus-last-frame image-to-video, combined image/video/audio references, and video editing. It generates synchronized audio-video clips lasting 4–15 seconds at 24 fps in 480p or 720p and supports several aspect ratios through an asynchronous task workflow. Its role is to retain frequently used multimodal creation functions in a lighter tier rather than match the standard model’s highest output specifications.

Key Capabilities

  • Lightweight Video Generation: Provides the Seedance 2.0 family’s foundational generation and editing workflows for everyday creation, concept validation, and frequent tasks.
  • Four-Modal Referencing: Combines text, images, video, and audio so separate assets can define subject identity, composition, style, movement, camera language, voice, and rhythm.
  • Text/Image-to-Video: Creates a scene from text or generates motion and transitions from a first frame or a first-and-last-frame image pair.
  • Native Synchronized Audio-Video: Produces speech, ambience, sound effects, and background music with the picture, with an option to generate silent video for separate audio post-production.
  • Reference-Based Editing: Uses image, video, and audio references to revise an existing clip, extending the workflow from one-pass generation to iterative editing.
  • Connected Clip Construction: Returns the generated result’s final frame for use as the first frame of another task, allowing multiple short generations to be organized into a connected sequence.

Technical Strengths

FeatureBenefit
Unified Multimodal Audio-Video DirectionExtends the Seedance 2.0 family’s joint-generation approach so picture, motion, dialogue, and effects can be coordinated in one creative task.
Complete Basic-Tier WorkflowCovers text, image, reference-based generation, and video editing in a lightweight tier, letting one model serve several creation stages.
Role-Based Multi-Asset InputsThe open platform accepts up to nine images, three video clips, and three audio clips, allowing creators to separate identity, action, camera, and sound references.
First-Frame and First/Last-Frame ControlAnchors the opening composition or defines both endpoints, supporting product transitions, logo motion, and short clips with a required landing frame.
Optional Synchronized AudioSwitches between a complete audiovisual prototype and silent visuals to accommodate different dubbing and post-production pipelines.
Asynchronous Tasks and Final-Frame ReuseSeparates submission, status checks, and result retrieval for batch management while reusing final frames to reduce discontinuity between clips.

Frequently Asked Questions

How does Mini differ from the standard and Fast variants?

Mini is the lightweight tier for broad production needs, Fast is positioned as the low-latency accelerated variant, and the standard model offers a higher output-specification ceiling. Public sources do not provide one standardized quality benchmark across the three, so teams should test their own identity-consistency, motion, text, and audio examples before choosing.

What prompt structure works well for batch generation?

Separate the prompt into fixed fields for subject, environment, action, camera, sound, and exclusions, changing only the variables for each batch. Repeat any required product color, logo, wardrobe, and final composition in every request instead of relying on previous tasks to preserve them.

How should first and last frames be used for a controlled transition?

The first frame should define the initial composition, subject state, and movement direction, while the last frame defines the end state and final composition. Describe the main action and camera path between them, avoiding new scenes or subjects that conflict with the required endpoint.

What if final delivery requires more than 720p?

Mini does not currently provide native 1080p or 4K output, so it is better used to validate composition, pacing, motion, and reference assets. For final delivery, verify a higher-resolution family variant on the access platform and regression-test the same prompt and references before switching.

What conditions apply when using a real person’s likeness?

ByteDance’s Seedance 2.0 notice says using a real portrait as a video subject reference requires identity verification or prior legal authorization. Users must also follow the platform’s safety rules, applicable personality-rights requirements, and local law.

Pricing

videosToken TypeLinkAI PriceOfficial Price
falseoutput$3.150000 / 1M tokens$3.500000 / 1M tokens
trueoutput$1.890000 / 1M tokens$2.100000 / 1M tokens

More from Bytedance