You just got access to MiniMax H3 and you're fired up: 2K output, image references, video references, audio-aware prompts — the whole multimodal toolbox. So you write the obvious prompt: a product, a style, a camera move, a mood, a color palette, an ending beat, all in one dense paragraph. The result comes back generic. The product morphs. The "mood" is everywhere and nowhere. Sound familiar?
The problem isn't effort. It's that H3 accepts four kinds of input — text, image, video, and audio — and nothing told you which input should own which part of the result. Cramming every intention into text is the fastest way to overload a prompt and get a muddy video.
This guide fixes that. It's current as of August 2026, checked against MiniMax's official H3 documentation (the model was announced on July 31, 2026). By the end, you'll know exactly when 2K is worth the cost, and how to assign one job to each input type so H3 works as a repeatable system instead of a lucky-dip generator.
Pain This Guide Solves
You want to use MiniMax H3's 2K resolution and its reference features, but you don't know which input type should control which part of the result. So you end up with overloaded prompts: one giant text block trying to carry identity, motion, mood, camera, and sound at the same time. That approach burns generations and produces videos where nothing is reliably controlled.
The fix is input control: give each job to the input that's best at it, keep the others quiet, and only reach for 2K when the direction already works.
What 2K Changes (And When Not to Use It)
Per MiniMax's official model documentation, H3 supports output up to 2K (768P to 2K), roughly 4–15 seconds at 24 fps. Resolution is the last thing you should spend budget on — it sharpens what's already good, but it can't invent direction.
Use 2K when:
- The clip is near-final and will be inspected closely (product surfaces, textures, logo detail)
- The output lands in a deck, ad, or landing page
- You need headroom for cropping or light editing in post
Skip 2K when:
- You're still exploring concepts and shots
- You're testing whether a motion idea even works
- You're iterating on prompt wording and will regenerate anyway
Rule of thumb: validate in 768P, commit in 2K. The resolution will still be there after the direction works.
One more timing note, so you're not surprised: H3 runs through the async API (you submit a task, then poll for the result), and as of August 2026 the paid video packages don't support H3 yet — there's no published price. Expect pay-per-use pricing to land with the platform rollout. MiniMax also says open weights are coming "in the coming days," so keep an eye on the release notes if you'd rather run the model yourself.
Text-to-Video: Best for Concept Exploration
Text is your idea input. It's the cheapest way to test a concept before you commit images, video, or resolution to it. Write a short production brief — subject, setting, action, camera, lighting, style boundary, ending beat:
A compact electric scooter stands under soft rain on a neon-lit street. The camera slowly pushes in as water beads roll across the black frame. Cool blue reflections, warm shop light in the background, realistic product-ad style, ending on the front wheel and logo.
Notice what's in that prompt: physical details that guide motion, a single scene, and one clear ending. It does not try to also define the exact scooter design, the actor, or the soundtrack. Those jobs go to other inputs.
If a text-only pass comes back with the right energy but the wrong subject, that's a signal — the next generation should hand the subject to an image.
Image-to-Video: Best for Product and Identity Control
When the subject matters — a product, character, costume, room, or brand asset — don't let text invent it. Upload a reference and tell H3 exactly what the image should control. The model accepts up to nine images per request (along with up to three videos and three audio clips, or twelve mixed files total), so you can be generous — but be specific about the job.
| Reference job | Prompt wording |
|---|---|
| Product shape and material | "Use the uploaded image for product shape and material." |
| Character identity | "Keep the face and hairstyle consistent with the reference." |
| Color palette | "Match the uploaded color palette, but change the background." |
| Composition | "Start from this framing and add a slow camera push." |
The rule of thumb: name the variable the image owns. "Use the reference" is a wish; "keep the product shape and material from the reference, change the scene" is an instruction. Ambiguous references produce ambiguous motion, because H3 doesn't know whether the image is about identity, style, composition, or color — you have to tell it.
Video-to-Video and Video Reference Workflows
A still image can't describe movement. That's the job of a video reference. Use one when the camera path, subject rhythm, or scene energy matters more than the pixels.
Good cases for a video reference:
- The camera path is the point (a dolly-in you want to replicate)
- You want the same action in a different setting
- You're testing an edit concept before committing
- A still image can't capture the motion (e.g., a walking cycle)
Keep the instruction narrow. "Match the walking rhythm and handheld camera feel" beats "make it like this video." The more precise the motion you name, the more control you keep over the result.
First and Last Frame Planning
For transitions and product reveals, think in frames. The first frame decides where the viewer starts; the last frame decides where the shot must land. When you plan both, H3 has a target instead of a vague "transition."
Reliable transition pairs:
- Product closed → product open
- Empty scene → full scene reveal
- Sketch → polished render
- Day mood → night mood
- Wide shot → close-up
Technical depth: MiniMax's H3 docs describe first-frame and last-frame as ways to pin the start and end states, and frame compatibility directly affects how clean the transition reads. The more the two frames share (angle, subject, lighting), the more usable the transition is likely to be. If you only have one strong frame, still upload it — but expect the weaker end of the transition to drift.
Competitor Difference: The Input Control Matrix
Most 2K guides stop at output quality — sharpness, upscaling, "tips for crispy results." That's table stakes. The part they skip is input control: deciding, before you generate, which input owns which job. Here's the matrix:
| What you need to control | Best input |
|---|---|
| Idea and mood | Text |
| Product identity | Image |
| Character consistency | Image |
| Motion style | Video reference |
| Start and end state | First and last frames |
| Sound cue | Audio-aware prompt (text input) |
Two notes on that last row. H3 is an omni-modal model with native audio, and it can take audio input — but audio can't be the sole prompt, and per the docs, sound is best steered by describing it in the text prompt (e.g., "rain on metal, distant traffic hum"). Uploaded audio clips work as extra conditioning, not as a replacement for a written brief.
Run your next generation through the matrix before you write anything: pick what you need to control, choose the owning input, and let everything else fall through. That one habit turns H3 from a novelty generator into a repeatable creative system — same inputs, predictable output.
A 2K-Ready 5-Step Workflow
Here's the whole system in one pass. Use it as your default H3 workflow:
- Concept in text. One scene, one action, one ending beat. No identity yet.
- Lock identity with an image. Add one reference and name what it owns ("product shape," "character identity").
- Fix motion with a video reference. Only if a still can't describe the movement.
- Plan the last frame. Describe where the shot lands before you render.
- Commit to 2K last. If the direction holds, rerun at higher resolution. If it doesn't, 2K will just render the weak idea sharper.
That sequence protects your budget and gives every generation a purpose.
Your First Low-Friction Test
Don't start with the scooter ad. Start smaller. Take one object you have a photo of — a mug, a shoe, a plant — and run this through MiniMax H3 AI:
"Use the uploaded image for shape and material. A slow 45-degree orbit around the object. Soft morning light, light dust in the air, subtle product-photo style."
Three inputs total: one image, one short prompt, no video, no 2K. You'll feel the difference immediately — identity comes from the image, motion from the text, and you can see exactly what each input contributed. If that works, promote the same structure to 2K. If it doesn't, adjust the prompt wording before you blame the resolution.
For larger jobs — real product ads, character-led clips, multi-scene concepts — come back to the matrix, then scale the workflow step by step. That's how you make H3 produce on a schedule instead of occasionally.
FAQ
Does MiniMax H3 support 2K?
Yes. Per MiniMax's official model documentation, H3 supports output up to 2K (768P/2K), roughly 4–15 seconds at 24 fps, accessed through the async API. Confirm current limits in the official docs before production — resolution tiers can change as the platform rolls out.
Should I use text, image, or video input?
It depends on the job. Text for ideas and mood, image for product identity and character consistency, video reference for motion style, and first/last frames for start and end states. Use the input control matrix above before each generation.
Why doesn't 2K make my video look better?
Resolution improves detail, not direction. If the prompt, reference, or camera instruction is weak, 2K renders the weak idea at higher resolution. Fix the input control first, then upgrade the resolution.
Can I use copyrighted references?
Only use references you own, assets you've licensed, or material you have permission to use. MiniMax's docs carry the same caution, and it applies to images, video clips, and audio alike. Review outputs before publishing.
When should I switch to 2K?
When the creative direction is validated: motion works, identity is consistent, and the clip is close to final. If you're still exploring concepts or iterating on prompts, stay at 768P. Reserve 2K for the near-final render — especially when the clip will be inspected closely in a deck, ad, or landing page.
Ready to put the matrix to work? Try a first test on the AI video generator page — one image, one short prompt, no 2K yet — and see how fast the system clicks.
Sources
- MiniMax H3 launch blog — Official H3 announcement and omni-modal model overview.
- MiniMax H3 release notes — Official H3 capability and rollout context.
- MiniMax video generation guide — Official video workflow guide.
- MiniMax create video generation task API — Official input and task parameters.
- MiniMax query video generation task API — Official result retrieval reference.
- MiniMax models introduction — Official model capability notes.
- MiniMax pay-as-you-go pricing — Official pricing source; verify current rates.
- MiniMax rate limits — Official rate limit source.
- MiniMax video packages — Official package support notes (H3 not yet supported as of Aug 2026).
- Hailuo AI website — Official consumer product context.





