Writing and testing prompts for MiniMax H3
Updated 2026-10-02
Most prompt advice for video models is folklore passed around as fact. This guide takes a different position: here is a structure that makes prompts easy to reason about, and here is a method for finding out what MiniMax H3 actually does with each part. VideoRouter's documentation does not describe H3-specific prompt behaviours, such as special keywords or camera syntax, so none are claimed here. What follows is a method to test, not a list of known tricks.
A four-part skeleton
Write prompts as four short clauses in a fixed order. A fixed order matters less for the model than for you: it lets you change one clause at a time and see what moved.
| Clause | Answers | Example |
|---|---|---|
| Subject | What is in the frame, concretely | a red paper airplane |
| Action | What changes over the clip | glides between two office towers, banks left |
| Camera | Where the viewer is and how that moves | wide shot, slow tracking behind the airplane |
| Style and light | Look, lighting, time of day | soft morning light, muted colour, film-like |
Joined, that is one sentence or two: "A red paper airplane glides between two office towers and banks left. Wide shot, slow tracking behind the airplane. Soft morning light, muted colour." Every clause is a claim you can verify by watching the output.
Principles that hold for any generator
- One action per clip. A five-second clip with four events asks the model to compress a storyboard. Split into several clips and cut them together, which also gives you control over each.
- Be concrete about the subject. "A person" leaves the model to choose; "a woman in a green raincoat holding a transparent umbrella" does not. Concrete nouns also make failures easy to see.
- Say what moves and what stays still. If the camera is supposed to be locked, say so. Ambiguity about motion is a common source of unwanted drift.
- Keep length proportional to duration. A long, dense prompt for a short clip invites partial execution.
- Prefer positive descriptions. Describe what you want to see. Whether H3 handles negations reliably is something to test, not assume.
These are working heuristics, not documented model rules. The next section is how you confirm or reject each on H3.
A one-variable-at-a-time test
Pick a base prompt and vary a single clause across a handful of values. Hold everything else fixed: model id, duration_secs, aspect_ratio, resolution. Because creation is an async job and polling is free, submitting all variants first and polling afterwards makes the whole grid cost about one clip's wall-clock time.
BASE = {
"subject": "a red paper airplane",
"action": "glides between two office towers and banks left",
"camera": "wide shot, slow tracking behind the airplane",
"style": "soft morning light, muted colour",
}
VARIANTS = {
"camera": [
"wide shot, locked-off camera",
"wide shot, slow tracking behind the airplane",
"low angle looking up, slow tilt",
],
}
def render(parts):
return f"A {parts['subject']} {parts['action']}. {parts['camera'].capitalize()}. {parts['style'].capitalize()}."
grid = []
for clause, values in VARIANTS.items():
for v in values:
parts = {**BASE, clause: v}
grid.append({"varied": clause, "value": v, "prompt": render(parts)})
# submit each with the same model/duration/aspect_ratio; store job ids with the grid row
Use the same grid on more than one subject, because a camera instruction that works on a paper airplane may do nothing on a portrait. Log the job id, the full prompt and the settings with every row. If you cannot reproduce which prompt made a clip, the test taught you nothing.
Scoring what you see
Write the rubric before you generate. Four questions are usually enough: Does the subject match the description? Does the motion match the action clause? Does the camera do what was asked? Are there visible artifacts such as warping hands or flicker? Score each from 0 to 2, have the people who will accept the output score it without seeing the prompt variant, and then join the scores back to the variants. A clause change is only a real effect if it shows up across several subjects, not one lucky clip.
Prompts for image-to-video
When you pass start_image_url, the image already supplies the subject and style, so the prompt's job shrinks to action and camera. Restating what the image shows is wasted words and can conflict with it. The image-to-video guide covers the request fields and failure modes.
Iterate cheaply, then spend
Run the grid on the cheapest H3 variant you can judge a composition on and the lowest tier that shows the effect you are testing. Then confirm the winners on the delivery tier, because a prompt finding on one tier may not carry over to another. Compare per-second cost across variants and hosts in the live table before you scale a grid, since the total is clips times seconds times rate. The model page at /models/ lists the hosts and tiers.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| minimax/h3 (768p) | MachGen $0.04 / second | WaveSpeedAI-resell $0.1 / second | 60% lower | 14 |
| minimax/h3-max (480p) | SandBase $0.01 / second | MiniMax $0.05 / second | 80% lower | 5 |
| minimax/h3-unrestricted (768p) | SandBase $0.064 / second | TOAPIS $0.08 / second | 20% lower | 2 |
Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
Once you know which clause changes matter for your content, freeze them into a template and version it next to your code. The Python tutorial has the create, poll and download scaffolding to run the grid, and videorouter.sh/signup gets you a key.
Frequently asked questions
Is there an official prompt syntax for MiniMax H3?
VideoRouter's documentation does not describe H3-specific prompt syntax or keywords. Use a plain structure of subject, action, camera and style, and test which parts change the output on your content.
How long should an H3 prompt be?
Keep it proportional to the clip length: one clear action per clip. Dense prompts for short clips risk partial execution, but test the length that works for your subjects.
Does the prompt matter for image-to-video?
Yes, but its job changes. The start image supplies subject and style, so the prompt should describe the motion and camera rather than restating the image.
How do I know whether a prompt change actually helped?
Change one clause at a time, keep model, duration, aspect ratio and resolution fixed, test on several subjects, and score the results blind against a rubric.
Keep reading
- MiniMax H3 Providers Compared — How to Pick a Host
- MiniMax H3 Image-to-Video API: start_image_url Guide
- MiniMax H3 Python Tutorial: Create, Poll, Download, Pin
- MiniMax H3 Resolution Tiers: How Per-Tier Pricing Works
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →