Minimax

MiniMax/MiniMax-H3

From $0.000000/ call

MiniMax's omni-modal audio-video generation system unifies text, image, video, and audio context to create videos up to 15 seconds at 2K with native stereo sound for advertising, e-commerce, and film production.

Text to VideoImage to Video

More from MiniMax

README

MiniMax/MiniMax-H3

MiniMax-H3 is a general-purpose omni-modal generation model released by MiniMax on July 31, 2026. It follows Hailuo 01 and Hailuo 02 as the H series' next step toward task generalization. The system jointly understands text, image, video, and audio context and generates video up to 15 seconds and 2K resolution with native 32 kHz stereo audio. H3-Base centers on a 33B-parameter dense, single-stream H3-Omni-Transformer, while the complete workflow also uses H3-Context-IR, H3-VisualVAE, H3-AudioVAE, and H3-Regenerate-2K.

Its key advance is to unify formerly separate tasks—including text-to-video, first/last-frame generation, subject, motion and style reference, audio conditioning, and audio-video editing—inside a natural-language-directed multimodal context. The same Transformer jointly predicts video and audio latents. For 2K output, the system regenerates the base result in context with the original inputs, giving it stronger evidence for reconstructing small text and fine textures than a standalone super-resolution stage. MiniMax has released two CFG-distilled H3-Base checkpoints; H3-Context-IR and H3-Regenerate-2K currently remain available through hosted APIs.

Key Capabilities

  • Unified Multimodal Understanding: Combines text with up to nine images, three video clips, and three audio clips, interpreting natural-language relationships between the references and the intended output.
  • Text-to-Audio-Video: Creates 4–15 second, 24 FPS videos from text while jointly producing native stereo tracks containing dialogue, sound effects, or music.
  • First/Last-Frame Control: Uses one image as a first or last frame, or two images to constrain both endpoints, giving creators direct control over a shot's opening and closing states.
  • Omni-Reference Generation: Draws on subjects or styles from images, motion or camera work from video, and cues from audio for cross-modal creation and video-to-video motion transfer.
  • Native Multi-Shot Modeling: Organizes multiple shots and their timing within one generation, supporting short advertisements, title sequences, and narrative clips.
  • 2K In-Context Regeneration: Reuses the 768p base output together with the original context to regenerate at 2K with better-grounded small text, brand elements, and fine detail.
  • Multilingual Dialogue: Reliably supports dialogue in 11 languages, including Chinese, English, Japanese, and Korean, for multilingual content production.

Technical Strengths

FeatureBenefit
Contextual Omni RepresentationUses language as a shared descriptive bridge across tasks and modalities, allowing complex reference relationships beyond fixed task templates.
33B Dense Single-Stream Omni TransformerProcesses and jointly predicts audio-video representations in one backbone, reducing modality silos and supporting broader task generalization.
High-Compression H3-VisualVAEUses 16× spatial and 4× temporal latent compression, followed by patchification for 32× effective spatial downsampling, reducing the cost of long visual sequences.
Dedicated Stereo AudioVAEEncodes and decodes left and right channels independently before recombination, preserving stereo spatial information during joint audio-video generation.
Three-Dimensional MM-RoPERepresents temporal, height, and width positions explicitly, helping maintain motion continuity and spatial structure.
In-Context RegenerationReuses the original multimodal conditions at the 2K stage, grounding details that conventional super-resolution often has to infer, including small text and fine textures.

Pricing

ResolutionLinkAI PriceOfficial PriceAdjustments
2K0.0000000.000000[object Object],[object Object],[object Object]
768P0.0000000.000000[object Object],[object Object],[object Object]