Generating many MiniMax H3 clips: a worker recipe
Updated 2026-10-02
Bulk video generation is mostly a bookkeeping problem. The model call is one POST; the work is making sure a crash, a rate limit or a flaky host does not lose jobs, double-submit them or run up an unbounded bill. This recipe builds that around the standard async POST /videos endpoint with minimax/h3 (or either variant), and starts by being clear about what VideoRouter's Batch API does and does not cover.
What the Batch API covers today
VideoRouter's Batch API (/v1/files and /v1/batches) is OpenAI-compatible and takes a JSONL file of requests. As documented, the only supported endpoint is /v1/chat/completions, so each line is a chat request. Video generation is not submitted through it, and there is no discounted batch tier for video. Do not plan an H3 workflow around a batch discount. The Batch documentation also describes a 15-minute flex mode and async jobs for chat, which are likewise chat features.
For video, bulk work therefore means many ordinary async jobs that you manage yourself. That is not a drawback, because creation returns immediately with a job id, polling is free, and the video endpoint already behaves like a queue.
Design: a manifest and a ledger
Keep two files. The manifest is the input: one JSON object per line with your own stable custom_id, the prompt and any fields such as start_image_url. The ledger is append-only output: one line per state change, keyed by custom_id. The ledger is what makes the run resumable and prevents duplicates, since the documented video request fields do not include an idempotency key. If your process dies, you read the ledger, skip anything that already has a job id, and continue.
The worker
import json, time, threading, requests
from concurrent.futures import ThreadPoolExecutor
API = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"}
LEDGER = "ledger.jsonl"
lock = threading.Lock()
def load_ledger():
seen = {}
try:
for line in open(LEDGER):
rec = json.loads(line)
seen[rec["custom_id"]] = {**seen.get(rec["custom_id"], {}), **rec}
except FileNotFoundError:
pass
return seen
def record(**rec):
with lock, open(LEDGER, "a") as f:
f.write(json.dumps(rec) + "\n")
def create(item, tries=4):
body = {"model": "minimax/h3", "duration_secs": 5, **item["fields"], "prompt": item["prompt"]}
for attempt in range(tries):
r = requests.post(f"{API}/videos", headers=H, json=body, timeout=30)
if r.status_code == 429:
time.sleep(int(r.headers.get("Retry-After", 5)))
elif r.status_code >= 500:
time.sleep(min(60, 2 ** attempt))
else:
r.raise_for_status() # 400/401/402/403: stop, do not retry
return r.json()["id"]
return None
def process(item, seen):
cid = item["custom_id"]
if seen.get(cid, {}).get("status") in ("completed", "failed"):
return
job_id = seen.get(cid, {}).get("job_id")
if job_id is None:
job_id = create(item)
if job_id is None:
record(custom_id=cid, status="create_failed")
return
record(custom_id=cid, job_id=job_id, status="submitted") # persist BEFORE polling
while True:
j = requests.get(f"{API}/videos/{job_id}", headers=H, timeout=30).json()
if j["status"] in ("completed", "failed"):
record(custom_id=cid, job_id=job_id, status=j["status"],
url=(j.get("data") or [{}])[0].get("url"), error=j.get("error"))
return
time.sleep(10)
items = [json.loads(l) for l in open("manifest.jsonl")]
seen = load_ledger()
with ThreadPoolExecutor(max_workers=4) as pool:
list(pool.map(lambda it: process(it, seen), items))
The one rule worth underlining is the comment: write the job id to the ledger before polling. That ordering is the difference between "resume after a crash" and "resubmit and pay twice". A create_failed record is deliberately not terminal in process, so a rerun will try it again.
Concurrency
Keys are rate limited per key with a token bucket, and a 429 carries a Retry-After header. Start with a small worker count, watch for 429s, and raise it only if you do not see them. More workers do not make an individual clip faster, and the bottleneck is usually generation time, not your submit rate. The same code also serves any other video model, so the worker count you settle on is a property of your key's limits, not of H3.
Retry rules for bulk
- 429: sleep for
Retry-Afterand retry creation. - 5xx at creation: every candidate host failed and nothing was billed, so exponential backoff is safe.
- Job
failed: a terminal state. Resubmit deliberately, ideally after reading the error. A content rejection will repeat. - Slow job: still running and still billed. Keep polling. Do not resubmit.
- Hedging: avoid
failover.on_timeout_secin bulk runs. It resubmits without cancelling the original and bills both attempts if both finish.
The Python tutorial covers single-job error handling in more detail.
Cost caps in layers
- Pre-flight estimate. Clips times seconds times the per-second rate for your tier, times
(1 + fee)with the platform fee a flat 2%. Refuse to start if it exceeds your run budget. Use live rates from the table below. - Per-run ceiling in code. Count submitted seconds in the worker and stop creating once the ceiling is reached.
- Per-key monthly cap. Set
monthly_spend_cap_usdon the key used for the run. Hitting it returns a 402 withspend_cap_exceeded, a hard stop that does not rely on your code being right. - Prepaid balance. A balance at or below zero returns 402 on live keys, so a run cannot spend more than you loaded.
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| minimax/h3 (768p) | MachGen $0.04 / second | WaveSpeedAI-resell $0.1 / second | 60% lower | 14 |
| minimax/h3-max (480p) | SandBase $0.01 / second | MiniMax $0.05 / second | 80% lower | 5 |
| minimax/h3-unrestricted (768p) | SandBase $0.064 / second | TOAPIS $0.08 / second | 20% lower | 2 |
Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
Run a ten-item pilot first and check actual billed seconds and attempts per accepted clip. If you are using the authentication docs' test keys for plumbing checks, read what they return before you rely on them. Then scale the manifest. The model page shows the host and tier grid, and videorouter.sh/signup gets you a key.
Frequently asked questions
Does the Batch API work for MiniMax H3 video?
No. As documented, the Batch API supports only the /v1/chat/completions endpoint. For video you submit async jobs to POST /videos and manage the queue yourself.
How do I avoid double-submitting jobs after a crash?
Keep an append-only ledger keyed by your own id and write the returned job id before you start polling. On restart, skip anything that already has a job id.
How many requests can I run in parallel?
That depends on your key's rate limits, which are enforced as a token bucket. Start small, watch for 429 responses with a Retry-After header, and raise concurrency only if none appear.
How do I stop a bulk run from overspending?
Estimate before starting, count submitted seconds in your worker, and set a monthly spend cap on the key. Reaching the cap returns a 402, and a zero prepaid balance does the same.
Keep reading
- MiniMax H3 Providers Compared — How to Pick a Host
- MiniMax H3 Image-to-Video API: start_image_url Guide
- MiniMax H3 Python Tutorial: Create, Poll, Download, Pin
- MiniMax H3 Resolution Tiers: How Per-Tier Pricing Works
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →