MiniMax launched H3 on July 31, 2026 — and if you are reading this in the first weeks of August, you are probably stuck in the gap between a launch post full of glossy demo clips and an API reference full of parameter tables. Can you download the weights? Is H3 the same thing as "Hailuo 3"? Can it really listen to audio as an input? And which of those demos can you actually reproduce on your own account? That gap is exactly what this guide exists to close.
It is an unusually messy moment to search "what is MiniMax H3": the model is days old, so most results are either press summaries of the same launch announcement or content-mill rewrites of it — rarely anything you could plan a project around. Everything below is checked against MiniMax's official H3 launch post and its platform documentation, both current as of August 2026. When a claim is future-dated or ambiguous, I say so instead of filling the gap with a guess.
After reading, you will be able to say precisely what H3 is and is not, choose the right generation mode for a job, name the constraints that will actually bite you, and produce a first test clip in minutes — through the raw API or through a browser tool like MiniMax H3 AI.
Pain This Guide Solves
Search for H3 and you will get one of two things: a launch post that shows off what the model can do, or a pricing page that tells you what it costs. Neither tells you what you can reliably depend on in a real workflow.
The pain points this guide addresses:
- The naming fog. "MiniMax H3," "Hailuo," "Hailuo 3," "MiniMax H3 AI" — four strings that describe three different layers, and they get used interchangeably.
- Spec vagueness. Launch demos are beautiful, but the useful numbers live in the API docs: resolution, duration, frame rate, input caps, file-size limits.
- The "open model" question. "Open weights coming in the coming days" gets paraphrased as "it's open source now." As of August 2026, that distinction matters.
- Mode confusion. H3 is not one prompt box. It has three distinct generation modes, and picking the wrong one wastes credits and queue time.
What This Guide Adds Over Typical Explainer Pages
Most SERP explainers for a new model follow the same three-act structure: repeat the press release, list features, and end with a generic "try it yourself." That structure serves the platform, not you. Here is what this guide does differently:
| Typical explainer | This guide |
|---|---|
| Calls everything "AI video generation" | Separates the model (MiniMax H3), the consumer product (Hailuo), and the wrapper tool (MiniMax H3 AI) |
| Lists features without numbers | Gives the official spec boundaries you can plan against |
| Quotes launch claims as settled facts | Flags what is confirmed vs. announced-but-not-yet-shipped (weights, technical report) |
| Assumes you know which mode to use | Gives a decision framework plus rule-of-thumb heuristics |
| No caveats | States the honest gotchas: closed-beta resolutions, package support, billing surprises |
If you already know the headline, skip straight to What MiniMax H3 Is for the substance, or to Try It First to test the model within minutes instead of reading about it.
What MiniMax H3 Is
MiniMax H3 is a general-purpose, omni-modal video generation model — the words matter. "Omni-modal" means it accepts text, images, video, and audio as inputs and understands them together, rather than only reading a typed prompt. It generates short video clips with native stereo sound output, up to 2K resolution, per MiniMax's official model documentation. MiniMax positions H3 as an "open" model: the launch post says model weights will be opened "in the coming days," subject to applicable laws and regulations, with a full technical report to follow. At the time of this writing (August 2026), the weights and the report had not yet been published, so treat "open source" claims you see elsewhere with caution.
That makes H3 a step beyond a plain text-to-video tool. A text-only model has to infer everything from words: the subject, the lighting, the camera move, the sound. H3 is built for a richer brief — a product image, a first frame, a last frame, a reference clip, an audio cue, or any combination of them, described in one natural-language sentence. MiniMax's launch material even shows this in its headline example: reference the camera movement from one video, have the character in a second image sing, with vocals matching a third audio clip. You describe the relationship; the model handles the blending.
The Three-Layer Model: H3 vs. Hailuo vs. MiniMax H3 AI
The single most useful thing you can learn from a "what is MiniMax H3" search is that the name you searched is a layer in a stack, and the layers are easy to confuse:
| Layer | Name | What it actually is |
|---|---|---|
| Model | MiniMax H3 | The generation model itself, accessed via the MiniMax API (model: "MiniMax-H3") |
| Consumer product | Hailuo | MiniMax's consumer-facing video app brand; "Hailuo 3" is informal shorthand for the H3 generation of that product |
| Wrapper / third-party tool | MiniMax H3 AI | A browser interface on top of the model, aimed at creators who do not want to wire the API themselves |
This distinction is not pedantry — it changes how you evaluate each layer. Model quality and spec limits come from the model documentation. Product behavior (queues, UI, upload UX) comes from the product experience. A wrapper tool is a separate decision, judged on workflow speed, upload support, pricing, and output handling rather than on the model's benchmarks. When you read a claim about "H3," ask which layer it is actually about before you act on it.
Which Generation Mode Should You Use?
H3's API exposes three generation scenarios, and they behave differently enough that choosing wrong costs you iterations. The official video generation guide defines them like this:
| Mode | What you give it | Best for | Gotcha |
|---|---|---|---|
| Text-to-Video | A text prompt only | Quick concepts, scratch ideas, style tests | Aspect ratio is required and cannot be "adaptive" |
| Image-to-Video (first/last frame) | Prompt + 1–2 images marked as first/last frame | Animating a specific image, or locking the exact opening and ending frames | Can't be combined with reference inputs — the two are mutually exclusive |
| Reference-to-Video | Prompt + any mix of reference images, videos, audio | Character/motion/style consistency, voice matching, V2V motion transfer | Needs at least one image or video; audio alone is rejected |
A practical rule of thumb: if you are testing a look, start with text-to-video. If you are testing consistency — a character, a product, a voice — jump straight to reference-to-video and give it a single strong reference image plus a short reference clip. And if you need a specific beginning and end, use first/last-frame rather than hoping the model guesses them.
This is also where a browser wrapper earns its keep: juggling roles like first_frame and reference_audio and re-uploading assets across attempts is exactly the friction a tool such as MiniMax H3 AI removes, since it manages uploads and mode selection for you while you focus on the prompt.
Spec Boundaries: Plan Against These Numbers
Every specific number below comes from MiniMax's official model documentation and API reference, checked August 2026. Treat them as the contract you plan against — not marketing:
| Constraint | Official limit | Planning implication |
|---|---|---|
| Output resolution | 768P or 2K | 2K costs more per second; at check time 768P was in closed beta (contact sales) |
| Output duration | 4–15 seconds, integers only | No continuous long renders — plan multi-shot stitching |
| Frame rate | 24 fps | Standard for short-form; fine for social and web |
| Reference images | Up to 9 | Generous — enough for a character sheet, not a dataset |
| Reference videos | Up to 3 clips, 2–15s each, 15s total | Reference clips are billed as input, so keep them short |
| Reference audio | Up to 3 clips, 2–15s each, 15s total; never alone | Audio must ride on an image or video reference |
| Mixed input cap | 12 files total | Budget your inputs before you assemble the call |
| Prompt length | 7,000 characters | Far more than you need — brevity still wins |
Two rules of thumb fall out of this table. First: H3 is a short-form model. If your idea needs a narrative arc, break it into 4–15-second shots and stitch them in an editor — that is how the use-case gallery was built, not by asking for one long clip. Second: audio is a companion, not a starter. You cannot prompt H3 with sound alone; audio informs the generation only when it is attached to an image or video reference.
A Technical Moment: The Async API, Straight from the Docs
MiniMax H3 video generation is asynchronous, and the flow is three steps: create a task, poll for status, then download the result. Per the official docs, creating a task returns a task_id; you poll with that ID until the status is succeeded, at which point the response includes content.url — the direct download link. Two details often catch people: query responses are only available for tasks from the last 7 days, and the download URL itself is time-limited, so fetch and store the file promptly rather than stashing the link.
A minimal end-to-end version, following the official guide's pattern:
import os, time, requests
api_key = os.environ["MINIMAX_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}"}
BASE = "https://api.minimax.io"
MODEL = "MiniMax-H3"
# Step 1: create the task
r = requests.post(f"{BASE}/v2/video_generation", headers=headers, json={
"model": MODEL,
"content": [
{"type": "text", "text": "A dancer on a drone doing flips, [zoom] camera"}
],
"duration": 5,
"resolution": "2K",
"ratio": "16:9",
})
task_id = r.json()["task_id"]
# Step 2: poll until succeeded
while True:
time.sleep(10)
t = requests.get(f"{BASE}/v2/query/video_generation/{task_id}",
headers=headers).json()["task"]
if t["status"] == "succeeded":
print(t["content"]["url"]) # Step 3: download this, then save locally
breakThe other technical detail worth understanding is the reference system, because it is H3's real differentiator. Inputs go into a content array where each item has a type (text, image_url, video_url, audio_url) and can carry a role (first_frame, last_frame, reference_image, reference_video, reference_audio). The roles define how the model reads your material. Two rules from the API reference: every request must include at least one non-empty text item, and image-to-video and reference-to-video inputs are mutually exclusive — mix first_frame with reference_video and the request fails. Under the hood, per the launch post, this multimodal understanding is powered by a "Contextual Omni Representation" that describes the relationships between your inputs, not just each input in isolation — which is why "make the character in Image 2 sing, matching the vocals in Audio 3" is a valid prompt.
Raw API or a Browser Tool? A Decision Framework
The API route is the right call when you are building a product: an app, a queue, an automated pipeline, or anything that needs repeatable, code-driven generation. The browser route is the right call when your goal is clips — testing prompts fast, uploading references, and downloading usable drafts without writing a single HTTP call. Both are legitimate; they serve different jobs.
| You want to... | Use |
|---|---|
| Build an app, queue, or automation around H3 | Raw MiniMax API |
| Generate clips from the browser, fast | MiniMax H3 AI |
| Test many prompt/ratio variants in an afternoon | MiniMax H3 AI (the API workflow is fine, just slower to iterate) |
| Own the full pipeline, billing, and retry logic | Raw MiniMax API |
If you are a creator, marketer, or small team, the browser route wins on loop speed: prompt → reference → generate → compare → revise, without touching request payloads. If you are an engineer, the raw API wins on control. There is no wrong answer as long as the decision is made on workflow, not on "which is more impressive."
Try It First: Your First H3 Clip in 3 Minutes
Before you commit to a workflow, verify the model actually does what you need with one cheap test. The lowest-friction version of this takes about three minutes:
- Open MiniMax H3 AI in a browser — no SDK setup, no auth plumbing.
- Enter a single, concrete prompt with an explicit camera move, e.g., "a ceramic mug on a wooden table, slow push-in, warm morning light."
- Set duration to 5 seconds and resolution to 768P (cheaper and plenty for a test).
- Generate, then check two things: motion quality on the camera move, and whether you get usable native audio back.
That one clip answers more questions than another hour of reading: it confirms the tooling works, gives you a real sense of generation speed, and surfaces any audio or motion gap before you invest in a bigger batch. If the clip matches your expectations, expand; if not, change your prompt style before you change your platform.
Limits, Honesty, and Caveats
Three honest caveats, because the launch hype will not tell you these:
- Open weights are promised, not shipped. As of August 2026, MiniMax has said weights open "in the coming days" and a technical report is coming, but neither had been released at check time. Follow MiniMax's official GitHub and Hugging Face orgs if you want to catch the release.
- Pricing pages lag the launch. At the August 2026 check, MiniMax listed per-second pay-as-you-go pricing for H3, but its video subscription/package pages did not yet list H3 support, and 768P was in a closed beta gated behind sales contact. Confirm current availability and numbers on the official pricing page before budgeting — do not rely on third-party summaries.
- It is still a production tool, not a magic button. Specs, prices, and supported modes will change; anything on this page is an August 2026 snapshot. Write prompts like shot briefs, use references you have rights to, and review outputs before they go into ads, product pages, or client work.
FAQ
Is MiniMax H3 the same as Hailuo 3?
No. MiniMax H3 is the model name; Hailuo is MiniMax's consumer-facing video product brand. "Hailuo 3" is informal shorthand people use for the H3 generation of that product, but the safest phrasing is always "MiniMax H3."
Can I download the MiniMax H3 weights today?
Not yet, as of August 2026. The official launch post says weights will open "in the coming days" and a technical report is "coming soon," but neither was published when this guide was written. Watch MiniMax's official GitHub and Hugging Face organizations if you want the release as soon as it lands.
Does MiniMax H3 generate audio?
Yes — native stereo audio output is part of the model, per MiniMax's materials, which is a meaningful step past silent-only video models. The nuance: audio cannot be your only input. In reference mode, audio must be accompanied by an image or video, and the generated soundtrack is baked into the output rather than handed to you as a separate stem, so plan for mixing in post if you need fine control.
How long can a MiniMax H3 video be?
Four to fifteen seconds, in integer values, per the official docs — at 768P or 2K and 24 fps. There is no "one long render" path, so longer content means planning multiple shots and stitching them in an editor.
What does MiniMax H3 cost?
MiniMax publishes per-second pay-as-you-go pricing for H3, and its launch post notes a price-performance advantage over mainstream models at both 2K and 768P. But at the August 2026 check, the video subscription/package pages did not yet list H3 support and 768P was in closed beta. Verify the current per-second rate and package availability on the official pricing page before you commit budget.
Bottom Line
MiniMax H3 is a multi-modal short-form video generation model that reads text, images, video, and audio together and outputs clips with native stereo sound — the real question is never "is it good" but "which mode, at which specs, through which interface." Plan around the official boundaries: 4–15 seconds, 768P or 2K, up to 9 reference images, audio always riding on a visual reference. And when in doubt, don't read about it — run one 5-second clip. The fastest way to do that today is to open MiniMax H3 AI, generate a single test, and judge the model on what comes back rather than on what was promised.
Sources
- MiniMax H3 launch post — Official July 31, 2026 announcement: positioning, native stereo sound, open weights "in the coming days," price-performance claims.
- MiniMax models release notes — Official release context placing H3 among MiniMax's model lineup.
- MiniMax video generation guide — Official workflow overview, generation modes, and input/output spec tables.
- MiniMax create video generation task API — Official task-creation parameters, content roles, and input limits.
- MiniMax query video generation task API — Official task status flow, 7-day query window, and download-URL handling.
- MiniMax models introduction — Official spec summary for MiniMax H3 (768P/2K, 4–15s, 24 fps).
- MiniMax pay-as-you-go pricing — Official pricing reference; verify current per-second rates and package support before purchase decisions.
- MiniMax rate limits — Official concurrency limits for MiniMax-H3 (Video Generation V2).
- MiniMax GitHub organization — Official org where H3 open-weight releases are expected to land.
- MiniMax Hugging Face organization — Official Hugging Face org for MiniMax model releases.





