MiniMax H3 API Guide: Video Generation Tutorial

Use the MiniMax H3 video API through LinkModel with verified model IDs, current pricing, multimodal inputs, asynchronous polling, and error handling.

MiniMax H3 API Guide: Video Generation Tutorial

TL;DR: Use MiniMax H3 through LinkModel by submitting model ID MiniMax-H3 to POST /v1/videos/generations, saving the returned task_id, and polling GET /v1/videos/generations/{task_id} until file_url is ready. H3 generates 4–15 second video at 768P or 2K with native stereo audio and optional image, video, or audio references. This guide uses LinkModel's current REST API rather than its still-supported historical routes.

What is the MiniMax H3 API?

MiniMax H3 is a general-purpose multimodal video model. It can interpret text together with images, video, and audio, then generate a new video with native stereo sound.

Through LinkModel, H3 uses the same asynchronous task pattern as other supported video models:

  1. Create a video-generation task.
  2. Receive a task_id immediately.
  3. Poll the query endpoint with backoff.
  4. Read file_url after the task succeeds.

What can MiniMax H3 generate?

The current MiniMax H3 model page exposes the following workflows:

WorkflowInputsTypical use
Text-to-videoPromptShort advertisements, title sequences, concept shots
First-frame videoPrompt and opening imageAnimate an existing composition
First-and-last-frame videoPrompt and two endpoint imagesControlled transitions
Reference-to-videoPrompt plus images, video, or audioIdentity, product, motion, style, or timing reference

H3 generates integer durations from 4 to 15 seconds. Available output resolutions are 768P and 2K. Text-to-video requires a fixed aspect ratio; requests with visual inputs can use adaptive.

How much does the MiniMax H3 API cost?

LinkModel's current H3 output price, checked August 27, 2026, is:

ResolutionPer output second5-second output10-second output15-second output
768P$0.08$0.40$0.80$1.20
2K$0.13$0.65$1.30$1.95

Reference inputs can change the final price. LinkModel's public H3 page currently lists $0.04 for each input image and bills reference-video seconds at the selected resolution rate. Review the current LinkModel runtime pricing before submitting a production workload. See the worked examples in the MiniMax H3 API pricing guide.

How do you use the MiniMax H3 API?

Step 1: create a text-to-video task

curl -X POST https://api.linkmodel.ai/v1/videos/generations \
  -H "Authorization: Bearer $LINKMODEL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-H3",
    "prompt": "Create a five-second product reveal: a glass bottle emerges from mist while the camera makes one slow clockwise orbit. Audio: quiet studio room tone and one soft glass chime.",
    "duration": 5,
    "resolution": "768P",
    "size": "16x9"
  }'

A successful create response uses LinkModel's standard envelope:

{
  "request_id": "req_abc123",
  "code": 0,
  "msg": "success",
  "data": {
    "task_id": "task_xyz789",
    "order_id": "order_abc123",
    "status": "Processing",
    "price": 0.4
  }
}

Store data.task_id. When present, order_id identifies the billing order and price is the charged amount. The generation continues after the create request returns. If this POST times out or its result is uncertain, do not blindly submit it again: the first request may already have created a task and billing order.

Step 2: poll the task status

curl "https://api.linkmodel.ai/v1/videos/generations/task_xyz789" \
  -H "Authorization: Bearer $LINKMODEL_API_KEY"

Status is Processing, Success, or Failed. On success, the response contains data.file_url:

{
  "request_id": "req_def456",
  "code": 0,
  "message": "success",
  "data": {
    "task_id": "task_xyz789",
    "status": "Success",
    "file_url": "https://cdn.your-domain.com/generated-video.mp4"
  }
}

The create response uses envelope field msg, while the task-query response uses message. Treat Success, Failed, and Cancelled as terminal states. Completed file_url values are dynamically signed, so download or copy assets you need instead of storing the URL as a permanent identifier.

Step 3: add polling backoff

import os
import time
import requests

BASE_URL = "https://api.linkmodel.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['LINKMODEL_API_KEY']}"}

create = requests.post(
    f"{BASE_URL}/videos/generations",
    headers={**HEADERS, "Content-Type": "application/json"},
    json={
        "model": "MiniMax-H3",
        "prompt": "A five-second cinematic product reveal with quiet room tone.",
        "duration": 5,
        "resolution": "768P",
        "size": "16x9",
    },
    timeout=30,
).json()

if create["code"] != 0:
    raise RuntimeError(f"create failed: {create['msg']}")

task_id = create["data"]["task_id"]
started = time.monotonic()
time.sleep(30)

while time.monotonic() - started < 900:
    poll = requests.get(
        f"{BASE_URL}/videos/generations/{task_id}",
        headers=HEADERS,
        timeout=30,
    ).json()

    if poll["code"] != 0:
        raise RuntimeError(f"poll failed: {poll.get('message', poll.get('msg'))}")

    status = poll["data"]["status"]
    if status == "Success":
        print(poll["data"]["file_url"])
        break
    if status in {"Failed", "Cancelled"}:
        raise RuntimeError(f"generation ended with {status}: {poll['data'].get('msg', '')}")

    elapsed = time.monotonic() - started
    time.sleep(5 if elapsed < 120 else 3 if elapsed < 240 else 8 if elapsed < 480 else 15)
else:
    raise TimeoutError("MiniMax H3 generation exceeded 15 minutes")

Persist the task ID before polling. If a worker restarts or a status request times out, resume polling the existing task rather than creating another potentially billable generation.

How do you use first and last frames?

Pass public HTTPS images as first_frame_image and last_frame_image:

{
  "model": "MiniMax-H3",
  "prompt": "Move from daytime to night while preserving the storefront, logo placement, and camera axis.",
  "first_frame_image": "https://cdn.your-domain.com/day.jpg",
  "last_frame_image": "https://cdn.your-domain.com/night.jpg",
  "duration": 8,
  "resolution": "2K",
  "size": "adaptive"
}

First/last-frame mode cannot be combined with reference-image, reference-video, or reference-audio inputs in the same request.

How do you add multimodal H3 references?

LinkModel exposes images, videos, audio_url, and audios for reference-to-video generation.

FieldCurrent limitWhat it should define
imagesUp to 9 URLsIdentity, product geometry, wardrobe, or style
videosUp to 3 URLsMotion, blocking, camera path, or timing
audiosUp to 3 URLsDialogue, voice, music, or rhythm

Reference audio requires at least one reference image or video. Public URLs must remain reachable while the task is created; LinkModel does not accept Base64 or data URIs for these H3 fields.

The prompt should assign one job to each input instead of merely attaching files:

Use Image 1 for the singer's identity and red suit.
Use Video 1 only for the slow handheld camera path.
Use Audio 1 for vocal timing and melody.
Do not copy people or background objects from Video 1.

See the MiniMax H3 video prompt guide for complete prompt templates.

Which MiniMax H3 parameters matter most?

LinkModel fieldAllowed valuePractical note
modelMiniMax-H3Case-sensitive public model ID
promptNon-empty stringRequired in every mode
durationInteger 4–15Default shown by LinkModel is 5
resolution768P, 2KDefault shown by LinkModel is 768P
size16x9, 1x1, 21x9, 3x4, 4x3, 9x16, adaptiveText-to-video needs a fixed ratio
first_frame_imagePublic URLCannot be mixed with reference mode
last_frame_imagePublic URLOptional partner to first frame
imagesUp to 9 public URLsReference mode
videosUp to 3 public URLsInput duration affects cost
audiosUp to 3 public URLsRequires an image or video reference

How should production apps handle H3 errors?

Separate transport failures from business errors. Model validation can return HTTP 200 with business code=400, so check both the HTTP status and response envelope:

  • Business code=400: inspect msg, correct the model or parameter, and create a new task.
  • HTTP 400: correct the malformed gateway request before retrying.
  • 401: replace the missing, invalid, or revoked API key; do not retry unchanged credentials.
  • 404 on task query: verify the path-based task ID.
  • 429: reduce concurrent generations and retry with backoff.
  • 500 or an uncertain POST timeout: check whether a task or billing order was created before submitting another generation.
  • Polling timeout: preserve the task ID for investigation instead of immediately creating a duplicate task.

Record the model, duration, resolution, reference package, estimated price, task ID, terminal status, output URL, and reviewer acceptance. These fields let you calculate cost per usable video rather than cost per submitted request.

Why use LinkModel for MiniMax H3?

LinkModel provides one API key, a consistent asynchronous task flow, and a shared model catalog for comparing H3 with other supported video models. Use the task-level price returned by the create response and run a small accepted-output test before moving production traffic.

Frequently asked questions

Is the MiniMax H3 API synchronous?

No. LinkModel creates an asynchronous H3 task and returns a task_id. Poll the video task endpoint until the status is Success, Failed, or Cancelled.

What is the LinkModel model ID for MiniMax H3?

Use the case-sensitive model ID MiniMax-H3. The MiniMax H3 model page shows the current schema, pricing rules, and Playground.

How much does the MiniMax H3 API cost?

LinkModel lists H3 output at $0.08 per second for 768P and $0.13 per second for 2K, before applicable reference inputs. Review the current runtime rules on the MiniMax H3 model page.

How do I use the MiniMax H3 API?

POST model MiniMax-H3 to /v1/videos/generations, capture the returned task ID, and poll /v1/videos/generations/{task_id} until status is Success, Failed, or Cancelled. On success, read the signed video URL from file_url and copy the asset you need to retain.

Does MiniMax H3 generate audio?

Yes. H3 jointly generates native stereo audio and video, and it can use reference audio when the request also includes an image or video reference.

About the author

Claire Lowe

Claire Lowe

AI and API researcher at LinkMode

Claire Lowe is an AI and API researcher at LinkModel, specializing in generative AI models, API pricing, provider comparisons, and multimodal infrastructure. Her work is grounded in official documentation, primary-source pricing data, and hands-on research, with a focus on helping developers and businesses make informed decisions about AI models and API providers.

Related Posts