MiniMax H3 Prompt Guide: How to Write Video Prompts That Work

MiniMax H3 Prompt Guide: How to Write Video Prompts That Work
Aug 12, 2026

MiniMax H3 Prompt Guide: How to Write Video Prompts That Work

Learn the official MiniMax H3 prompt format, camera and audio syntax, reference-mode structure, and a low-cost testing loop for video prompts.

You've watched the MiniMax H3 demo clips: 2K output, native stereo sound, a camera that actually listens, and voices that lip-sync. Then you write your first prompt the way you always have — "cinematic drone shot of a city, golden hour, epic music" — and the result comes back with a camera that drifts randomly, no sound at all, and a style that has nothing to do with "cinematic."

That gap — between what the demo promised and what your first clip delivered — is not bad luck. It's a prompt-format mismatch.

MiniMax H3 is not an image model that happens to make video. It builds a full audiovisual timeline from your text: shots, cuts, camera moves, who speaks when, ambient sound, and the music track are all things the prompt can — and per MiniMax's own guidance, should — control explicitly. If you keep writing prompts the way you did for image generation or earlier video models, H3 leaves most of your intent on the table.

This guide is current as of August 12, 2026, and it is built directly on MiniMax's official public documentation: the company's own Video Prompt Writing Guide for H3 (published on the official MiniMaxAI Hugging Face repository in early August 2026), the official H3 launch post, and the official API documentation. Where something comes from the official guide, you'll see the exact syntax. Where something is a judgment call (like how to test cheaply), you'll see it labeled as such.

By the end, you'll be able to write an H3 prompt that controls the shot sequence, the camera, the dialogue, and the soundscape — and you'll know exactly what to change when the output ignores you.

The Pain This Guide Solves

The core problem for most H3 beginners is not a lack of ideas — it's that the prompt has an implicit structure you've never been told about, and without it you can't diagnose failures.

You probably do one of two things today: you write a short moody line ("a sad girl by the window, rain") and accept whatever comes back, or you write a giant comma-stuffed paragraph covering style, subject, action, camera, sound, and mood all at once. The first approach leaves H3 with almost no direction; the second overloads it so that no single instruction is followed well.

Neither approach tells you why a failure happened. Did H3 ignore the camera move because the camera wasn't specified in the right format? Because the shot was cut mid-move? Because the audio intent was buried under 300 words of style? Without a structure, every failed generation is a coin flip, and since H3 bills per second, coin flips get expensive.

Competitive Difference: The Official Format, Not a Collection of Tips

Most MiniMax H3 prompt guides circulating right now are opinionated tip lists — "describe the scene first," "direct the camera," "use cuts intentionally." They're not wrong, but they're someone's summary of the official material, often written before MiniMax published its full prompt-writing guide.

MiniMax's official H3 repository on Hugging Face (MiniMaxAI/MiniMax-H3) includes the Video Prompt Writing Guide: a specification for how an H3 prompt should be structured, with a three-field body, exact camera-motion syntax, speaker IDs, dialogue markup, shot timestamps, and complete worked examples for text-to-video, image-to-video, first/last-frame, and reference mode. That document is the ground truth for this article. You get the same format MiniMax's own documentation uses, plus the decision framework and troubleshooting that the spec itself doesn't spell out.

Step 1: Understand How H3 Reads Your Prompt

Before any syntax, you need the mental model: an H3 prompt is a script, not a description. The official guide describes the main prompt field as building a "complete audiovisual timeline from text" — visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline, plus a summary of ambient sound and a separate field for background music.

That's why one-sentence prompts underperform. A script has parts. H3's official structure has three core fields:

integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
  • integrated_multimodal_description — the main body. Style, composition, subjects, actions, shot changes, dialogue, and any sound that happens inside the scene.
  • overall_soundscape — a 1–4 sentence summary of ambient and physical sound across the whole video (wind, rain, footsteps, impacts).
  • non_diegetic_music — background music only the audience hears, described in 1–3 sentences by instrumentation, tempo, and dynamics — not by mood words.

You don't need to reproduce the official rewriting pipeline to benefit from it. The practical takeaway: when you plan a prompt, plan three things — the action timeline, the ambient sound, and the music — and write them as separate blocks. One practical note from the official API docs: the text prompt has a hard limit of 7,000 characters per request, so there's room to be explicit without rushing.

Step 2: Write the Main Description Like a Script

The integrated_multimodal_description is where 80% of your prompt's power lives. Four sub-systems matter most.

2.1 Shots and Cuts

The official guide writes shots as [Shot 1], [Shot 2], and so on. The first shot carries no timestamp; each later shot begins with a strictly increasing cut time inside the video duration:

[Shot 1] Live-action, cinematic, a medium-wide shot frames...
[Shot 2] At 00:03.500, the camera cuts to...

Cut phrases are plain English: "the camera cuts to," "the shot transitions to," "the shot changes to." Rule of thumb from the official guide: a cut should introduce genuinely new information — subject, space, state, viewpoint, or time. If only the distance or a slight angle changes, use camera motion instead of a cut. Two shots at 5 seconds with nothing new between them is how you get two murky half-scenes.

2.2 Camera Motion: Type + Amplitude + Speed

Camera direction is the most commonly attempted — and most commonly ignored — part of H3 prompts. The official guide's syntax is a three-part expression: motion type, amplitude, speed, written as natural English inside the shot rather than a tag at the end of the sentence.

Motion types the official guide defines: Zoom In/Out, Push In/Pull Out, Pan Left/Right, Truck Left/Right, Tilt Up/Down, Pedestal Up/Down, Arc Shot, Tracking Shot, Static Shot, Shake Slightly/Strongly, POV, and Roll Clockwise/Counterclockwise. Amplitude is with small amplitude or with large amplitude; speed is at slow speed or at fast speed. Medium amplitude and normal speed can be omitted.

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.

Notice what's absent: "dolly," "tracking" used as a generic word, or camera tags stacked at the end. The API docs also mention a shorthand — bracketed labels like [pan], [zoom], [static] directly after key descriptions — which is a lighter-weight way to steer the camera in simple prompts. Both are official; the full sentence form is the one the official writing guide uses, and it gives the model more to work with.

2.3 Speakers, Dialogue, and On-Screen Text

If anyone speaks, sings, or produces an off-screen voice, the official guide assigns a stable ID per vocal source: (S1), (S2), and (S1,S2) for group speech. The ID stays with the same speaker across all shots. When a speaker first appears, anchor their identity — character type, age, pitch, timbre, speaking rate — before the line. Dialogue itself goes inside <d> with a language tag, and the words must be preserved verbatim:

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

For voiceover, the official guide is precise: use the phrase "says in an off-screen voiceover," and immediately state that the character's lips remain closed:

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

If a line crosses a cut, mark the connection with <scenetrans> at the join and state that the audio continues across the cut; use <cutoff> when speech is truncated by the end of the video. And any text visibly on screen — signs, labels, subtitles, neon — goes in English double quotation marks, verbatim:

A red neon sign reading "营业中" glows above the doorway.

2.4 Soundscape and Music Fields

The two audio fields are where H3's "native stereo sound" actually gets directed:

  • overall_soundscape — ambient and physical sound in one continuous paragraph: wind, rain, traffic, footsteps, fabric, impacts, breathing. Dialogue and singing stay out of this field (they belong in the main description). Use N/A only when you want complete silence.
  • non_diegetic_music — audience-only score described by instrumentation, speed, rhythm, and dynamics. Not mood words ("epic," "sad"), which the official guide explicitly avoids.
overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

This is also how you get lip-sync and synchronized sound to work: audio intent is a first-class citizen of the prompt, not an afterthought.

Step 3: Add Image Alignment for I2VA, FL2VA, and L2VA

When you add a reference image, the official guide prefixes the prompt with an alignment instruction that says exactly where the image lands in the timeline. This is the instruction that prevents "the model ignored my reference image."

For image-to-video (the image is the opening frame):

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

For first-and-last-frame:

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

For last-frame-only (L2VA), only the endpoint is anchored, and the description must "infer a plausible earlier state" that converges onto the image by the final shot. In all three cases the alignment line comes first, followed by a blank line, then the three core fields.

The description style differs by mode too. For I2VA, anchor the shot on the image's style, subject, composition, and scene before describing the next action — "first-frame anchor → action onset → continuous development → result." For FL2VA, don't describe the two static images; describe the motion path between them, and favor a single shot so the model can interpolate continuously.

Step 4: Reference Mode Uses a Six-Section Structure

Character references, motion references, and voice references are H3's headline feature — the launch post itself demonstrated "reference the camera movement from Video 1, have the character in Image 2 sing, with vocals matching Audio 3." When you build a prompt around reference assets, MiniMax's official full-reference guide writes the prompt in six sections:

subject_definitions     — what each referenced asset means and its label
summary                 — task type + target video in one paragraph
retention_analysis      — what is preserved, transferred, or referenced
detailed_description    — the shot-by-shot timeline (normally 350–500 words)
overall_soundscape      — ambient and physical sound
non_diegetic_music      — audience-only score

Labels are the load-bearing part. subject_definitions assigns four label types: <Subject N> for reusable content (a person, a scene, a style), <Picture N> for images that anchor a frame, <Video N> for whole-video structure (editing, continuation, rhythm), and <Audio N> for audio that is copied or referenced. A voice reference becomes a definition like:

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

Then detailed_description uses those labels where they apply in each shot, and retention_analysis marks each label's relationship: fully_preserved, partially_preserved, attribute_transfer, weak_reference (for visuals) and fully_copy, partially_copy, reference, weak_reference (for audio). That last section is your diagnostic: if a reference asset is being modified when you wanted it preserved, the fix is usually in how the label was defined, not in how you described the shot.

You don't have to hand-build this six-section structure for every generation — H3's official API includes an H3-Context-IR endpoint that interprets your multimodal context and returns a structured, enhanced prompt for you. But understanding the structure is how you judge whether the enhanced prompt (or your own) is any good.

Step 5: Test Cheap, Then Spend

Because H3 is billed per second of output, the testing loop is a cost problem, not just a quality problem. Three rules keep iterations affordable:

  1. Test at 768P and short duration. Direction, camera, and audio intent can all be validated at 768P in 4–5 seconds. Resolution is a finishing feature; per MiniMax's pricing docs, 2K costs more per second, so don't pay for it while you're still debugging whether the prompt works. (If 768P isn't available in your tool, note that MiniMax's docs describe it as a closed beta — fall back to testing with shorter 2K clips.)
  2. Change one variable per generation. Pick the single weakest link — the camera move, the dialogue, the soundscape — and rewrite only that. Prompt files with three simultaneous changes tell you nothing about which one worked.
  3. Keep a written log. Attempt → change → outcome, three or four lines each. After three or four passes you'll have a direction that works; only then raise resolution or length for the real shot.

This is also where the three-field structure pays off: a failed camera move means you edit the camera sentence, not the whole prompt.

Troubleshooting: Why H3 Ignores Your Prompt

SymptomLikely causeFix
Output ignores the camera moveCamera written as a tag stack at the end, or amplitude/speed missingWrite type + amplitude + speed as a natural sentence inside the shot
Reference image has no effectNo alignment instruction, or image buried without a labelPrefix the alignment line; in reference mode, define the asset in subject_definitions
No sound in the outputAudio intent only implied, or stuck in mood wordsAdd overall_soundscape; describe the music by instrumentation and tempo in non_diegetic_music
Lips move with no voice (or voice with no lips)Voiceover rules not followedUse "says in an off-screen voiceover" and state the lips remain closed
Everything drifts mid-videoToo many actions per shot, or cuts without new informationOne action per shot; cut only when the scene genuinely changes, otherwise use camera motion
Character changes identity between shotsNo stable speaker/subject ID, no reference lockKeep (S1) across shots; use a reference image defined as <Subject N>
Clip is too short or too longNon-integer durationUse an integer between 4 and 15 seconds
Prompt feels generic despite detailDescription is a list of attributes, not a timelineAdd [Shot N] structure with sequential actions and cut times

FAQ

What is the best prompt format for MiniMax H3?

The official structure from MiniMax's Video Prompt Writing Guide: an alignment instruction when images are used, then three fields — integrated_multimodal_description (the shot-by-shot timeline), overall_soundscape (ambient sound), and non_diegetic_music (audience-only music). Reference mode expands to six sections including subject_definitions and retention_analysis.

How long should a MiniMax H3 prompt be?

The official API limit is 7,000 characters per text prompt. There's no official minimum, but the official guide's worked examples show that a structured multi-shot description with soundscape and music — roughly 100–300 words — is a normal working length. Short mood lines usually underperform because they leave the timeline underspecified.

Can MiniMax H3 generate sound from the prompt?

Yes. H3 generates video with native stereo sound, and the prompt controls it: dialogue and diegetic sounds in the main description, ambient sound in overall_soundscape, and background music in non_diegetic_music. Describing music by instrumentation and tempo works far better than mood words.

How do I keep a character consistent across shots?

Give the speaker a stable ID (S1) that never changes between shots, and for visual identity, provide a reference image. In reference mode, define the character in subject_definitions (for example, "the woman whose appearance comes from Picture 1") and reuse the same label in every shot's description.

Does MiniMax H3 support prompts in Chinese or other languages?

The prompt itself works in multiple languages, and dialogue inside <d> preserves the original language verbatim — for example <d>[中文] 我在下一站下车。</d> with a language tag. Per MiniMax's official writing guides, the structured fields themselves are expected in English (the rewriting guides are written in English), which keeps the shot and camera syntax unambiguous.

Is there a tool that writes the prompt for me?

MiniMax's official API includes H3-Context-IR, a task type that interprets your text, image, video, and audio context and returns a structured, enhanced prompt. It doesn't generate video — it produces the prompt, which you can then feed into a generation task. Understanding the format in this guide is still useful: you'll be able to tell whether the enhanced prompt is actually good.

How much does it cost to test prompts?

H3 is billed per second of output. Testing at 768P with short durations is the cheapest path, and 2K carries a higher per-second rate per MiniMax's pay-as-you-go pricing page — verify current numbers there since pricing can change.

Core Summary

MiniMax H3 treats your prompt as a script for a full audiovisual timeline, not a wish list. Write the main description with explicit shots, camera moves (type + amplitude + speed), speaker IDs, and verbatim dialogue; direct ambient sound and music in their own fields; anchor reference images with an alignment instruction; and test every change cheaply at short duration before spending on resolution and length.

The fastest way to feel the difference between a described prompt and a scripted prompt is to generate the same idea both ways. If you'd rather start generating than read more theory, try your first structured H3 prompt in MiniMax H3 AI — it handles resolution, duration, and references for you, and new accounts get free credits to run the testing loop above without spending a dollar.

For the full workflow around reviewing and iterating your generations, see How to Use MiniMax H3; for when 2K is worth the cost, see the MiniMax H3 2K Text, Image, and Video Guide; and for building this into an automated pipeline, see the MiniMax H3 API Guide.

Sources

Accuracy note (August 12, 2026): this article is compiled from the official sources above. MiniMax H3 launched July 31, 2026, and its documentation is evolving quickly; the syntax in the official writing guides may be updated, so treat this guide as a structured reading of the official material rather than a replacement for it. Where this article offers judgment (testing strategy, troubleshooting), it is labeled as such and is not official MiniMax guidance. We do not claim hands-on testing of every syntax element described.

Essayez le générateur vidéo AI Minimax H3 gratuit

Transformez une invite, une photo ou un clip de référence en 2K natif avec son, via le modèle Minimax H3 AI. Crédits gratuits à l'inscription, sans carte.