Home › Guides › Image-to-video

Animating an image with the MiniMax H3 API

Updated 2026-10-02

Image-to-video with MiniMax H3 uses the same endpoint and model id as text-to-video. What changes is one request field and a handful of decisions about hosts, prompts and resolution. This guide covers those, plus the reference-to-video option that goes beyond a single start frame.

The minimal request

import requests

resp = requests.post(
    "https://videorouter.sh/api/v1/videos",
    headers={"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"},
    json={
        "model": "minimax/h3",
        "prompt": "the subject turns toward the camera and smiles, gentle push-in",
        "start_image_url": "https://example.com/portrait.jpg",
        "duration_secs": 5,
        "resolution": "768p",
        "aspect_ratio": "16:9",
    },
)
job = resp.json()
print(job["id"], job["status"])

start_image_url accepts a public https:// URL or an inline data:image/...;base64,... URI. Omit it and the same model id performs text-to-video. The response is a job you poll at GET /videos/{id} until completed or failed. Polling is free, and the cost is charged once at creation from the requested duration. The Python tutorial has the full poll-and-download loop.

Where the image can come from

Host support

H3 is served by many hosts, and image input is wired on most but not all of them. The video docs name Fal, Atlas Cloud, Replicate, WaveSpeedAI and MachGen as hosts where H3 accepts start_image_url. The model pages list every host with its per-resolution prices and show the accepted input fields. Two practical points:

Reference-to-video: more than one frame

If a single start frame is not enough to carry a character or product, H3 has a reference mode on some hosts. Per the docs, minimax/h3/fal and minimax/h3/atlas-cloud can composite up to 9 images, 3 videos and 3 audio clips (12 combined) in one call. You pass them as arrays:

json = {
    "model": "minimax/h3/fal",
    "prompt": "the character from the reference image walks through the market",
    "input_references": [
        {"type": "image_url", "image_url": {"url": "https://example.com/character.jpg"}}
    ],
    "input_video_references": [
        {"type": "video_url", "video_url": {"url": "https://example.com/motion.mp4"}}
    ],
    "duration_secs": 5,
}

At least one reference across the three arrays is required, and exceeding the per-kind or combined cap is rejected with a 400. The mode is chosen automatically by which fields you send, so there is no separate model id. MachGen's H3 reference row is narrower: images only, up to 7. The video docs also state that end_image_url (last-frame control) is not accepted by any model yet.

Prompts for a still

The image already defines subject, framing and light. Spend the prompt on what changes:

  1. Camera: static, push-in, slow pan, orbit.
  2. Subject action: one clear verb, such as turns, lifts, walks, smiles.
  3. Ambient motion: wind, steam, water, light shifts.
  4. Guardrails: what should stay fixed, such as the label or the face.

Describing the whole scene again tends to compete with the image. Keep it to a sentence or two, and change one variable between tests so you learn something from each run. Results are prompt-dependent, so build a small evaluation set of your own images instead of relying on generic examples.

Choosing a resolution for image-to-video

Each host lists a different set of tiers, and the price spread between hosts changes by tier. Two rules are enough for most work:

Duration is the other cost driver. You pay for the requested seconds, so ask for the shortest clip that demonstrates the motion during testing.

Common failure modes

SymptomLikely causeFix
400 on createHost or model does not accept the field you sentPick a host that supports image input, or remove the unsupported field
Job failedImage unreachable, unsupported, or rejected by content policyCheck the URL from outside your network; read job["error"]
Output is the wrong sizeResolution or ratio not supported, so the default was usedCheck the model page for supported tiers
402Balance empty or key cap reachedTop up or raise the cap

Jobs that fail upstream are not billed. See the quickstart to make a first call, and create a key at videorouter.sh/signup.

Frequently asked questions

How do I do image-to-video with MiniMax H3?

Send start_image_url alongside the prompt to POST /videos with model: minimax/h3. Omit the field and the same model does text-to-video.

Which hosts accept an image for H3?

Per the docs, Fal, Atlas Cloud, Replicate, WaveSpeedAI and MachGen. The model page lists every host and the input fields it accepts.

Can I use several reference images with H3?

On minimax/h3/fal and minimax/h3/atlas-cloud, yes: up to 9 images, 3 videos and 3 audio clips (12 combined) via input_references and the video and audio reference arrays.

Can I control the final frame?

Not currently. end_image_url is documented as accepted by no model yet and is rejected with a 400.

Keep reading

Using MiniMax H3 is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →