Home › Guides › H3 vs H3 Max test harness

A reusable harness for comparing minimax/h3 and minimax/h3-max on your prompts

Updated 2026-10-02

Public opinion about which variant is better is not a substitute for your prompts. This page gives you a harness that runs the same prompt set through minimax/h3 and minimax/h3-max, records enough to reconstruct every run, and a rubric for human review. It makes no claims about which model wins; its job is to make your own comparison repeatable. The Python tutorial has a client you can reuse for the plumbing.

Design rules for a fair comparison

The prompt set

Use 10 to 20 prompts drawn from real work, grouped by what you want to stress: camera motion, human subjects, text or fine detail, lighting, and your single hardest case. Store them in a JSON file with an id and a category so results can be sliced later.

[
  {"id": "p01", "category": "camera", "prompt": "slow dolly through a market at dusk"},
  {"id": "p02", "category": "people", "prompt": "two people talking at a cafe table, handheld"}
]

The harness

import csv, json, os, time, requests
from concurrent.futures import ThreadPoolExecutor

API = "https://videorouter.sh/api/v1"
H = {"Authorization": f"Bearer {os.environ['VIDEOROUTER_KEY']}",
     "Content-Type": "application/json"}
MODELS = ["minimax/h3", "minimax/h3-max"]
PARAMS = {"duration_secs": 5, "resolution": "768p", "aspect_ratio": "16:9"}
REPEATS = 2
MAX_PARALLEL = 3          # stay well under your key's rate limit

def run_one(task):
    prompt_id, category, prompt, model, rep = task
    base = dict(prompt_id=prompt_id, category=category, model=model, rep=rep)
    t0 = time.time()
    r = requests.post(f"{API}/videos", headers=H, timeout=30,
                      json={"model": model, "prompt": prompt, **PARAMS})
    if not r.ok:
        try:
            msg = r.json().get("error", {}).get("message", "")
        except ValueError:
            msg = r.text[:200]
        return dict(base, job_id="", status=f"http_{r.status_code}",
                    seconds="", url="", error=msg)
    job = r.json()
    while job["status"] not in ("completed", "failed"):
        if time.time() - t0 > 1500:
            break
        time.sleep(5)
        job = requests.get(f"{API}/videos/{job['id']}", headers=H, timeout=30).json()
    return dict(base, job_id=job["id"], status=job["status"],
                seconds=round(time.time() - t0, 1),
                url=(job.get("data") or [{}])[0].get("url", ""),
                error=str(job.get("error") or ""))

prompts = json.load(open("prompts.json"))
tasks = [(p["id"], p["category"], p["prompt"], m, rep)
         for p in prompts for m in MODELS for rep in range(REPEATS)]

with ThreadPoolExecutor(MAX_PARALLEL) as ex:
    rows = list(ex.map(run_one, tasks))

with open("ab_results.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=rows[0].keys())
    w.writeheader(); w.writerows(rows)

Latency here is wall-clock time from submission to a terminal status, which includes queueing. It is a property of that moment and that host, not a stable number for either model, so repeat at different times of day before reading anything into it.

Recording cost without hard-coding prices

Do not type rates into the script. Compute expected cost from the live table for the tier and host you pinned, multiplied by the requested duration, and reconcile against your balance or usage records afterwards. Billing is once at creation from the requested duration, polling is free, and jobs that fail upstream are not billed, so failed rows should not count toward spend. The live comparison:

ModelCheapest hostPriciest hostCheapest isHosts
minimax/h3 (768p)MachGen
$0.04 / second
WaveSpeedAI-resell
$0.1 / second
60% lower14
minimax/h3-max (480p)SandBase
$0.01 / second
MiniMax
$0.05 / second
80% lower5
minimax/h3-unrestricted (768p)SandBase
$0.064 / second
TOAPIS
$0.08 / second
20% lower2

Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

Add an expected_cost column that you fill from that table or from your own price file, and treat it as an estimate until checked against actual usage.

Blind review

Write a second script that downloads each url, renames the file to a random token, and stores the token-to-job-id mapping privately. Reviewers see only the tokens. Present the two clips for one prompt side by side, in random left/right order.

import csv, json, os, uuid, requests

os.makedirs("review", exist_ok=True)
rows = [r for r in csv.DictReader(open("ab_results.csv")) if r["status"] == "completed"]
key = {}
for r in rows:
    token = uuid.uuid4().hex[:8]
    open(f"review/{token}.mp4", "wb").write(requests.get(r["url"]).content)
    key[token] = r["job_id"]
json.dump(key, open("blind_key.json", "w"))   # keep this away from reviewers

Copy outputs to your own storage promptly rather than treating result URLs as a permanent archive.

A review rubric

Ask each reviewer to score each clip from 1 to 5 on the following, then choose a preferred clip per pair:

DimensionQuestion
Prompt adherenceDoes the clip contain what the prompt asked for?
MotionIs movement coherent, free of warping or jitter?
Subject stabilityDo faces, hands and objects keep their shape across frames?
DetailIs fine detail clean at the size you will deliver?
UsabilityWould you ship it, fix it in post, or regenerate?

Aggregate by category, not only overall. A pattern such as "no difference except on people" is more actionable than a single winner, because you can route by category in production. Compare preference counts to the price difference on your pinned host before deciding.

What the results can and cannot tell you

Route by what you learned and keep the harness for the next model release. See the model pages for hosts and tiers, the pricing page for the live comparison, and create a key to run it.

Frequently asked questions

How do I compare minimax/h3 and minimax/h3-max fairly?

Send identical prompts and parameters to both, pin the same host so you compare models rather than hosts, repeat important prompts, and have reviewers rank clips blind.

How many prompts do I need?

Ten to twenty representative prompts is a practical start, including your hardest case, with two or three repeats of the prompts that matter most because generation is not deterministic.

Does the harness tell me which model is better?

No. It records job ids, status, latency and outputs so that human reviewers can compare. The conclusion applies to your prompts, one host and one point in time.

Are failed test jobs billed?

Per VideoRouter's documentation, jobs that fail upstream are not billed. Successful jobs are billed once at creation from the requested duration, and polling is free.

Keep reading

Using MiniMax H3 is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →