Animating an image with the MiniMax H3 API
Updated 2026-10-02
Image-to-video with MiniMax H3 uses the same endpoint and model id as text-to-video. What changes is one request field and a handful of decisions about hosts, prompts and resolution. This guide covers those, plus the reference-to-video option that goes beyond a single start frame.
The minimal request
import requests
resp = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
json={
"model": "minimax/h3",
"prompt": "the subject turns toward the camera and smiles, gentle push-in",
"start_image_url": "https://example.com/portrait.jpg",
"duration_secs": 5,
"resolution": "768p",
"aspect_ratio": "16:9",
},
)
job = resp.json()
print(job["id"], job["status"])
start_image_url accepts a public https:// URL or an inline data:image/...;base64,... URI. Omit it and the same model id performs text-to-video. The response is a job you poll at GET /videos/{id} until completed or failed. Polling is free, and the cost is charged once at creation from the requested duration. The Python tutorial has the full poll-and-download loop.
Where the image can come from
- A hosted URL. It must be reachable by the provider, so signed URLs with short expiry and private-network addresses will fail.
- A data URI. Fine for small images. Large base64 payloads inflate your request.
- The uploads endpoint.
POST /v1/uploadstakes a multipart file up to 50 MB and returns a presigned URL valid for 30 minutes. Use it immediately in the next call. It is scratch space, not storage.
Host support
H3 is served by many hosts, and image input is wired on most but not all of them. The video docs name Fal, Atlas Cloud, Replicate, WaveSpeedAI and MachGen as hosts where H3 accepts start_image_url. The model pages list every host with its per-resolution prices and show the accepted input fields. Two practical points:
- Leave routing automatic if you only care about price. If a host cannot take your input, expect a 400 rather than a silent text-to-video fallback, and move to a host that can.
- If you want a specific host, append it:
minimax/h3/fal. That is a soft preference. For a hard pin use"provider": {"only": ["fal"], "allow_fallbacks": false}.
Reference-to-video: more than one frame
If a single start frame is not enough to carry a character or product, H3 has a reference mode on some hosts. Per the docs, minimax/h3/fal and minimax/h3/atlas-cloud can composite up to 9 images, 3 videos and 3 audio clips (12 combined) in one call. You pass them as arrays:
json = {
"model": "minimax/h3/fal",
"prompt": "the character from the reference image walks through the market",
"input_references": [
{"type": "image_url", "image_url": {"url": "https://example.com/character.jpg"}}
],
"input_video_references": [
{"type": "video_url", "video_url": {"url": "https://example.com/motion.mp4"}}
],
"duration_secs": 5,
}
At least one reference across the three arrays is required, and exceeding the per-kind or combined cap is rejected with a 400. The mode is chosen automatically by which fields you send, so there is no separate model id. MachGen's H3 reference row is narrower: images only, up to 7. The video docs also state that end_image_url (last-frame control) is not accepted by any model yet.
Prompts for a still
The image already defines subject, framing and light. Spend the prompt on what changes:
- Camera: static, push-in, slow pan, orbit.
- Subject action: one clear verb, such as turns, lifts, walks, smiles.
- Ambient motion: wind, steam, water, light shifts.
- Guardrails: what should stay fixed, such as the label or the face.
Describing the whole scene again tends to compete with the image. Keep it to a sentence or two, and change one variable between tests so you learn something from each run. Results are prompt-dependent, so build a small evaluation set of your own images instead of relying on generic examples.
Choosing a resolution for image-to-video
Each host lists a different set of tiers, and the price spread between hosts changes by tier. Two rules are enough for most work:
- Iterate low, deliver high. Judge motion on a lower tier, then re-run the approved shots at the tier you ship. The tier guide explains how to read the per-tier table.
- Match the aspect ratio to the source. Supported ratios per the docs are 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9, but a model supports only some combinations, and an unsupported
resolutionoraspect_ratiois ignored rather than rejected. Check the returned file's dimensions withffprobein your test suite.
Duration is the other cost driver. You pay for the requested seconds, so ask for the shortest clip that demonstrates the motion during testing.
Common failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| 400 on create | Host or model does not accept the field you sent | Pick a host that supports image input, or remove the unsupported field |
Job failed | Image unreachable, unsupported, or rejected by content policy | Check the URL from outside your network; read job["error"] |
| Output is the wrong size | Resolution or ratio not supported, so the default was used | Check the model page for supported tiers |
| 402 | Balance empty or key cap reached | Top up or raise the cap |
Jobs that fail upstream are not billed. See the quickstart to make a first call, and create a key at videorouter.sh/signup.
Frequently asked questions
How do I do image-to-video with MiniMax H3?
Send start_image_url alongside the prompt to POST /videos with model: minimax/h3. Omit the field and the same model does text-to-video.
Which hosts accept an image for H3?
Per the docs, Fal, Atlas Cloud, Replicate, WaveSpeedAI and MachGen. The model page lists every host and the input fields it accepts.
Can I use several reference images with H3?
On minimax/h3/fal and minimax/h3/atlas-cloud, yes: up to 9 images, 3 videos and 3 audio clips (12 combined) via input_references and the video and audio reference arrays.
Can I control the final frame?
Not currently. end_image_url is documented as accepted by no model yet and is rejected with a 400.
Keep reading
- MiniMax H3 Providers Compared — How to Pick a Host
- MiniMax H3 Python Tutorial: Create, Poll, Download, Pin
- MiniMax H3 Resolution Tiers: How Per-Tier Pricing Works
- MiniMax H3 Prompt Guide: A Testable Structure for Video Prompts
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →