Home › Guides › Batch generation

Generating many MiniMax H3 clips: a worker recipe

Updated 2026-10-02

Bulk video generation is mostly a bookkeeping problem. The model call is one POST; the work is making sure a crash, a rate limit or a flaky host does not lose jobs, double-submit them or run up an unbounded bill. This recipe builds that around the standard async POST /videos endpoint with minimax/h3 (or either variant), and starts by being clear about what VideoRouter's Batch API does and does not cover.

What the Batch API covers today

VideoRouter's Batch API (/v1/files and /v1/batches) is OpenAI-compatible and takes a JSONL file of requests. As documented, the only supported endpoint is /v1/chat/completions, so each line is a chat request. Video generation is not submitted through it, and there is no discounted batch tier for video. Do not plan an H3 workflow around a batch discount. The Batch documentation also describes a 15-minute flex mode and async jobs for chat, which are likewise chat features.

For video, bulk work therefore means many ordinary async jobs that you manage yourself. That is not a drawback, because creation returns immediately with a job id, polling is free, and the video endpoint already behaves like a queue.

Design: a manifest and a ledger

Keep two files. The manifest is the input: one JSON object per line with your own stable custom_id, the prompt and any fields such as start_image_url. The ledger is append-only output: one line per state change, keyed by custom_id. The ledger is what makes the run resumable and prevents duplicates, since the documented video request fields do not include an idempotency key. If your process dies, you read the ledger, skip anything that already has a job id, and continue.

The worker

import json, time, threading, requests
from concurrent.futures import ThreadPoolExecutor

API = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"}
LEDGER = "ledger.jsonl"
lock = threading.Lock()

def load_ledger():
    seen = {}
    try:
        for line in open(LEDGER):
            rec = json.loads(line)
            seen[rec["custom_id"]] = {**seen.get(rec["custom_id"], {}), **rec}
    except FileNotFoundError:
        pass
    return seen

def record(**rec):
    with lock, open(LEDGER, "a") as f:
        f.write(json.dumps(rec) + "\n")

def create(item, tries=4):
    body = {"model": "minimax/h3", "duration_secs": 5, **item["fields"], "prompt": item["prompt"]}
    for attempt in range(tries):
        r = requests.post(f"{API}/videos", headers=H, json=body, timeout=30)
        if r.status_code == 429:
            time.sleep(int(r.headers.get("Retry-After", 5)))
        elif r.status_code >= 500:
            time.sleep(min(60, 2 ** attempt))
        else:
            r.raise_for_status()           # 400/401/402/403: stop, do not retry
            return r.json()["id"]
    return None

def process(item, seen):
    cid = item["custom_id"]
    if seen.get(cid, {}).get("status") in ("completed", "failed"):
        return
    job_id = seen.get(cid, {}).get("job_id")
    if job_id is None:
        job_id = create(item)
        if job_id is None:
            record(custom_id=cid, status="create_failed")
            return
        record(custom_id=cid, job_id=job_id, status="submitted")   # persist BEFORE polling
    while True:
        j = requests.get(f"{API}/videos/{job_id}", headers=H, timeout=30).json()
        if j["status"] in ("completed", "failed"):
            record(custom_id=cid, job_id=job_id, status=j["status"],
                   url=(j.get("data") or [{}])[0].get("url"), error=j.get("error"))
            return
        time.sleep(10)

items = [json.loads(l) for l in open("manifest.jsonl")]
seen = load_ledger()
with ThreadPoolExecutor(max_workers=4) as pool:
    list(pool.map(lambda it: process(it, seen), items))

The one rule worth underlining is the comment: write the job id to the ledger before polling. That ordering is the difference between "resume after a crash" and "resubmit and pay twice". A create_failed record is deliberately not terminal in process, so a rerun will try it again.

Concurrency

Keys are rate limited per key with a token bucket, and a 429 carries a Retry-After header. Start with a small worker count, watch for 429s, and raise it only if you do not see them. More workers do not make an individual clip faster, and the bottleneck is usually generation time, not your submit rate. The same code also serves any other video model, so the worker count you settle on is a property of your key's limits, not of H3.

Retry rules for bulk

The Python tutorial covers single-job error handling in more detail.

Cost caps in layers

  1. Pre-flight estimate. Clips times seconds times the per-second rate for your tier, times (1 + fee) with the platform fee a flat 2%. Refuse to start if it exceeds your run budget. Use live rates from the table below.
  2. Per-run ceiling in code. Count submitted seconds in the worker and stop creating once the ceiling is reached.
  3. Per-key monthly cap. Set monthly_spend_cap_usd on the key used for the run. Hitting it returns a 402 with spend_cap_exceeded, a hard stop that does not rely on your code being right.
  4. Prepaid balance. A balance at or below zero returns 402 on live keys, so a run cannot spend more than you loaded.
ModelCheapest hostPriciest hostCheapest isHosts
minimax/h3 (768p)MachGen
$0.04 / second
WaveSpeedAI-resell
$0.1 / second
60% lower14
minimax/h3-max (480p)SandBase
$0.01 / second
MiniMax
$0.05 / second
80% lower5
minimax/h3-unrestricted (768p)SandBase
$0.064 / second
TOAPIS
$0.08 / second
20% lower2

Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

Run a ten-item pilot first and check actual billed seconds and attempts per accepted clip. If you are using the authentication docs' test keys for plumbing checks, read what they return before you rely on them. Then scale the manifest. The model page shows the host and tier grid, and videorouter.sh/signup gets you a key.

Frequently asked questions

Does the Batch API work for MiniMax H3 video?

No. As documented, the Batch API supports only the /v1/chat/completions endpoint. For video you submit async jobs to POST /videos and manage the queue yourself.

How do I avoid double-submitting jobs after a crash?

Keep an append-only ledger keyed by your own id and write the returned job id before you start polling. On restart, skip anything that already has a job id.

How many requests can I run in parallel?

That depends on your key's rate limits, which are enforced as a token bucket. Start small, watch for 429 responses with a Retry-After header, and raise concurrency only if none appear.

How do I stop a bulk run from overspending?

Estimate before starting, count submitted seconds in your worker, and set a monthly spend cap on the key. Reaching the cap returns a 402, and a zero prepaid balance does the same.

Keep reading

Using MiniMax H3 is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →