← Back to Blog
MiniMaxPrompt EngineeringVideo GenerationGuides

MiniMax H3 Video Prompt Guide: Examples

Learn MiniMax H3 prompting with the official three-part formula, multimodal reference roles, shot-by-shot examples, audio direction, and API templates.

2026-08-27

MiniMax H3 Video Prompt Guide: Examples

TL;DR: MiniMax's official H3 formula is reference asset instructions + core concept + shot-by-shot description. Assign every uploaded file a role with labels such as @Image 1, describe the target video in one sentence, then direct each shot's framing, visible action, camera movement, dialogue, and sound. The examples below turn that formula into reusable templates for MiniMax H3, rather than presenting an unofficial benchmark.

What makes a good MiniMax H3 video prompt?

Use the three-part structure published in the latest MiniMax H3 Cookbook:

Complete prompt = Reference asset instructions + Core concept + Shot-by-shot description

A practical production checklist still helps inside those three sections: goal, subject and setting, timeline, camera, audio, and constraints. It expands the official formula; it does not replace it.

For a reference-driven request, begin by assigning one job to every input in upload order:

Use @Image 1 for the character's identity and wardrobe.
Use @Video 1 only for body motion and camera movement.
Use @Audio 1 for vocal timing, melody, and the exact lyrics quoted below.

This opening block prevents H3 from guessing whether a file defines identity, composition, motion, style, voice, timing, or the entire soundtrack. Also state what must not transfer from a reference, such as the actor or background in a motion-only video.

How should you structure a MiniMax H3 prompt?

1. Write reference asset instructions

Use MiniMax's upload-order labels and define the intended contribution of every file. Common roles include:

AssetUseful roles to assignDetails to preserve or exclude
@Image 1Character, product, environment, style, composition, opening or ending keyframeIdentity, geometry, wardrobe, logo, text, or camera axis
@Video 1Character motion, blocking, camera motion, timing, performance, or editing patternSay whether its people, objects, wardrobe, and background must be ignored
@Audio 1Voice, dialogue, lyrics, rhythm, emotion, full soundtrack, or a selected time rangeQuote exact words when speech accuracy matters

An audio reference cannot be used by itself; pair it with at least one image or video reference.

2. State the core concept in one sentence

Define the subject, location, action, genre or style, camera treatment, movement, and editing approach:

Create a 10-second premium skincare film in a clinical white studio, using controlled macro photography, slow lateral camera movement, and three clean cuts from texture to product to hero frame.

Avoid metaphors that cannot be photographed. “Premium” becomes useful only after the prompt makes its visible meaning concrete through lighting, materials, framing, movement, and pace.

3. Direct the video shot by shot

Break the clip down by timeline or story progression when it needs multiple shots. For each segment, specify framing, visible content, camera movement, character action, dialogue, and sound:

0–3s — Macro close-up: condensation gathers on the bottle. The camera trucks slowly right. Quiet studio room tone.
3–8s — Medium product shot: the bottle rotates counterclockwise while the camera trucks left and pans right to orbit it. A clean light sweep reveals the label.
8–10s — Centered hero shot: camera locked off. One soft glass chime; no dialogue.

H3 generates 4–15 second clips and uses cuts by default. Ask for a continuous take only when that is the intended treatment; do not combine “one continuous shot” with a list of incompatible cuts.

4. Make camera direction executable

Separate camera motion from subject motion and name the shot size, viewpoint, focus, and movement. For example, an orbit is clearer as truck left + pan right or truck right + pan left than as “orbiting camera.”

Camera: medium shot at product height, truck left while panning right, shallow focus held on the label.
Subject: the bottle rotates counterclockwise at a constant speed.

This makes it clear whether the camera, subject, or both should move.

5. Treat audio as part of the direction

H3 jointly generates native stereo audio, so describe sound as part of the scene:

Audio: quiet studio room tone, a soft glass chime at the reveal, no dialogue, no music.

For dialogue, identify the speaker as on-screen, off-screen, or voice-over; quote the exact line; and leave enough screen time for it to be spoken naturally. H3 supports J-cuts and L-cuts, so state when dialogue begins before a cut or continues after it. MiniMax reports reliable TTS coverage for 11 languages and exploratory support for more than 40 others, but pronunciation and lip synchronization still need review before publication.

If no score is wanted, write Non-diegetic music: N/A or explicitly request no additional background music. Do not ask for both music and no BGM.

6. Protect the non-negotiables

End with a short constraint list:

Keep the bottle geometry, cap color, and label spelling unchanged. No extra products or hands.

Use concrete, non-conflicting restrictions. Protect the details that would make the output unusable without contradicting the shot plan.

What are useful MiniMax H3 prompt examples?

Text-to-video product advertisement

Create a 10-second 16:9 premium product film for a clear glass perfume bottle on black stone.
0–3s: a narrow spotlight reveals the silhouette through drifting mist.
3–7s: the camera makes a slow clockwise orbit while the bottle remains still.
7–10s: gold reflections travel across the glass and settle into a centered hero frame.
Audio: low room tone, one restrained glass chime at 7s, no voice.
Keep the bottle proportions and label text stable. No hands, duplicate bottles, or camera shake.

With no references, describe appearance, environment, motion, and sound more fully. A useful sequence moves from a wide establishing view to a medium action shot and then a close-up detail.

First-and-last-frame transition

Transform the daytime storefront in the first frame into the illuminated nighttime storefront in the last frame.
The camera performs one slow forward push. Pedestrians cross naturally and lights turn on progressively from inside to outside.
Audio: distant city traffic and soft footsteps, no music.
Preserve the storefront architecture, logo placement, and camera axis throughout the transition.

Use first_frame_image and last_frame_image for this mode. H3 does not allow first/last-frame inputs to be mixed with reference-image, reference-video, or reference-audio mode in the same request.

Multimodal character performance

Create a 12-second single-shot performance on a small jazz-club stage.
Use @Image 1 for the singer's identity and red suit. Use @Video 1 only for the slow handheld camera path. Use @Audio 1 for vocal timing and melody.
The singer looks toward camera for the final line, then gives a restrained smile.
Preserve facial identity, suit details, microphone position, and the spatial relationship between performer and band.
Do not copy people, wardrobe, or background objects from @Video 1.

The key is explicit reference ownership: identity from one input, motion from another, and timing from audio.

Precise product-video edit

Reference asset instructions:
Use @Video 1 as the source video. Replace only the bottle with the product in @Image 1. Preserve the original hand motion, timing, camera movement, background, lighting direction, reflections, and audio.
 
Core concept:
Create a seamless product replacement that looks captured in the original shot.
 
Shot-by-shot description:
0–4s: preserve the hand reaching into frame and wrapping around the bottle. Match the replacement bottle to the original perspective and occlusion.
4–8s: preserve the lift and camera tracking. Keep the @Image 1 label facing camera without changing its spelling.
8–10s: preserve the final hero pose and existing sound. Do not alter the hand, sleeve, table, or background.

For editing prompts, identify the target change and explicitly protect everything else. H3's official examples include object and character replacement, wardrobe and background changes, relighting, dialogue replacement, local visual effects, and other selective edits.

Vertical creator video

Create a 9:16 handheld morning-routine clip in a small sunlit bathroom.
0–4s: the creator places the phone near the mirror and applies moisturizer.
4–8s: autofocus briefly shifts to the product, then returns to the face.
8–12s: the creator smiles and steps out of frame.
Audio: natural room echo, fabric movement, and one casual spoken sentence in English: "This is my five-minute routine."
Keep the handheld movement subtle and physically plausible. No beauty-filter skin or floating objects.

Small imperfections such as autofocus shifts should be requested only when they support the intended creator aesthetic.

Which MiniMax H3 generation mode should you use?

The official Cookbook groups H3 workflows into three practical modes:

ModeStart withPrompt emphasis
Multimodal asset fusionCharacter, product, motion, and audio referencesAssign a distinct role to every asset and resolve any overlap
Image-to-videoOne image, or opening and ending framesDefine whether the image is the opening, ending, identity, or environment reference
Text-to-videoA text brief without media referencesDescribe appearance, environment, movement, shot progression, and sound in greater detail

With two keyframes, H3 creates the connecting motion, lighting, and sound rather than automatically inserting a cut. In multimodal mode, an audio reference must be accompanied by an image or video.

How do image, video, and audio references work in H3?

The official H3 API supports up to 9 reference images, 3 reference videos, and 3 reference audio files, with no more than 12 media files total. Each audio clip can be 2–15 seconds and audio references can total up to 15 seconds; audio also requires an image or video reference. Reference videos can total up to 15 seconds. A prompt remains required even when references are supplied.

Use this reference map before writing the final prompt:

InputAssign one primary jobAvoid
Reference imageIdentity, product geometry, wardrobe, or styleAsking one image to define incompatible angles
Reference videoMotion, blocking, camera path, or timingAccidentally importing its subject or background
Reference audioDialogue timing, voice, music, or rhythmLeaving the desired speaker relationship unclear

Every input needs an explicit purpose. If two references appear to control the same feature, clarify which one has priority.

When should you use the MiniMax Hailuo 2.3 structure?

For short visual-only MiniMax Hailuo 2.3 jobs, LinkModel uses Camera–Subject–Reaction as a compact editorial shorthand. It is not the official H3 formula:

Camera: slow push-in at eye level.
Subject: an anime courier in a yellow raincoat walks through a neon alley.
Reaction: hair and coat sway in the wind; puddles ripple under each step; she notices the camera and smiles.
Style: hand-painted anime background, controlled line work, warm face against cool city light.

Use it for one subject, one dominant action, and one visual result. Move to H3's fuller structure when you need native audio, multiple timed beats, cross-modal references, first/last frames, or 2K delivery.

How can you diagnose a weak MiniMax video result?

SymptomLikely prompt problemRevision
H3 ignores uploaded assetsTheir purpose is not definedAssign each file a role with @Image, @Video, or @Audio labels
Shots feel incoherentThe prompt is one unstructured paragraphSplit it into reference instructions, core concept, and shot-by-shot direction
Correct scene, wrong motionAction is vagueName actor, direction, speed, and end state
Camera and subject both driftTheir motions are conflatedPut camera and subject on separate lines
Reference identity changesInput roles are unclearState which input owns identity and what must remain fixed
Audio feels genericSound is described as mood onlySpecify source, timing, environment, and spoken words
Audio direction conflictsThe prompt asks for music and no BGMChoose one instruction and state it consistently
Cuts appear in a continuous takeShot structure contradicts camera directionChoose either a continuous take or an explicit edit sequence
Final seconds rushToo many beatsReduce events or assign a short timeline
Product details mutateNo protected attributesList geometry, colors, and text that must remain unchanged

If camera motion is wrong, revise the camera line before replacing every style adjective. When identity is essential, supply a character reference instead of expecting a short text description to hold a face across cuts.

How do you send an H3 prompt through the API?

LinkModel exposes H3 as an asynchronous video task. This minimal text-to-video request uses the verified public model ID and fields:

curl -X POST https://api.linkmodel.ai/v1/videos/generations \
  -H "Authorization: Bearer $LINKMODEL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-H3",
    "prompt": "Create a five-second product reveal: a glass bottle emerges from mist while the camera makes one slow clockwise orbit. Audio: quiet room tone and one soft glass chime.",
    "duration": 5,
    "resolution": "768P",
    "size": "16x9"
  }'

The response returns a task_id. Poll GET /v1/videos/generations/{task_id} until the status is Success, Failed, or Cancelled; on success, read the dynamically signed file_url and copy the asset you need to retain. The complete create, poll, reference-input, and error-handling flow is in the MiniMax API guide.

Frequently asked questions

What is the best MiniMax H3 prompt structure?

Use MiniMax's official three-part structure: reference asset instructions, a one-sentence core concept, and a shot-by-shot description. Within it, define the goal, setting, timeline, camera, audio, constraints, and the job of every uploaded asset.

Can MiniMax H3 use image, video, and audio references together?

Yes. H3 can combine reference images, videos, and audio, but audio cannot be used alone and must be paired with at least one image or video. Assign a clear job to every input and explain its relationship to the target video.

What is the difference between H3 text-to-video, image-to-video, and multimodal generation?

Text-to-video requires the most detailed description of appearance, movement, shots, and sound. Image-to-video anchors the opening, ending, identity, or environment, while multimodal generation combines separate references for elements such as identity, motion, camera behavior, voice, and rhythm.

How do you keep a character consistent in MiniMax H3?

Provide a clear character image, assign it the identity role, and list the facial, hair, and wardrobe details that must remain unchanged. After a cut, identify the returning character explicitly instead of assuming H3 will infer continuity from the previous shot.

How do you control camera movement in a MiniMax H3 prompt?

Write camera motion separately from subject motion and use executable directions such as shot size, viewpoint, truck, pan, tilt, push-in, handheld movement, and focus behavior. For an orbit, specify truck left + pan right or the reverse rather than only saying orbiting camera.

Does MiniMax H3 generate dialogue and sound with video?

Yes. H3 generates native stereo audio with the video; prompts can specify exact dialogue, the speaker, delivery, ambience, sound effects, and music. MiniMax reports reliable TTS coverage for 11 languages, but dialogue timing, pronunciation, and lip synchronization should still be reviewed.

Can MiniMax H3 edit an existing video?

Yes. Use the source video as a reference, name the exact character, object, wardrobe, background, lighting, dialogue, or effect to change, and explicitly preserve everything else. This targeted instruction reduces unintended changes to motion, timing, framing, and audio.

How long can a MiniMax H3 video and prompt be?

The H3 API accepts video durations from 4 to 15 seconds and prompts up to 7,000 characters. A longer prompt is useful only when it clearly assigns references and directs multiple shots; it should not add conflicting camera, edit, or audio instructions.

Why does MiniMax H3 ignore a reference image, video, or audio file?

The most common prompt-level cause is an uploaded asset with no defined role or two assets competing to control the same feature. Label inputs in upload order, state what each one contributes, set a priority when roles overlap, and identify details that must not transfer.

Should a MiniMax H3 prompt include negative instructions?

Include concrete, non-conflicting constraints that protect identity, geometry, exact text, or continuity. Make sure those restrictions do not contradict the shot plan or audio direction; a restriction that conflicts with the requested shots gives the model no single target to follow.

Where can I test MiniMax H3 prompts and API requests?

Use the MiniMax H3 model page to review current parameters and run a prompt in the Playground. When moving to production, follow the MiniMax API guide for task creation, status polling, terminal-state handling, and retrieval of the signed output URL.

From brief to shot

Test a structured H3 prompt

Choose a duration and resolution, run the prompt in the Playground, then inspect continuity, audio, and protected details.

Sources: the latest MiniMax H3 official Cookbook, MiniMax H3 video prompt guide, LinkModel First API Call, LinkModel MiniMax H3 API reference, and the MiniMax H3 announcement. Last checked August 27, 2026.

Related Posts