A reusable harness for comparing minimax/h3 and minimax/h3-max on your prompts
Updated 2026-10-02
Public opinion about which variant is better is not a substitute for your prompts. This page gives you a harness that runs the same prompt set through minimax/h3 and minimax/h3-max, records enough to reconstruct every run, and a rubric for human review. It makes no claims about which model wins; its job is to make your own comparison repeatable. The Python tutorial has a client you can reuse for the plumbing.
Design rules for a fair comparison
- Same inputs. Identical prompt,
duration_secs,resolutionandaspect_ratiofor both models. Change only themodelstring. - Same host. Hosts can differ in behaviour. To compare models rather than hosts, hard-pin one host for both with
provider.onlyandallow_fallbacks: false, provided both models are on it per the model pages. Amodel/hostsuffix alone is a soft preference and can still fall back. - Short clips. You pay once at creation from the requested duration, so keep test durations small.
- Repeat the hard prompts. Generation is not deterministic, so one sample per prompt says little. Run two or three per cell for the prompts you care about most.
- Blind review. Reviewers should not see which model made which clip.
The prompt set
Use 10 to 20 prompts drawn from real work, grouped by what you want to stress: camera motion, human subjects, text or fine detail, lighting, and your single hardest case. Store them in a JSON file with an id and a category so results can be sliced later.
[
{"id": "p01", "category": "camera", "prompt": "slow dolly through a market at dusk"},
{"id": "p02", "category": "people", "prompt": "two people talking at a cafe table, handheld"}
]
The harness
import csv, json, os, time, requests
from concurrent.futures import ThreadPoolExecutor
API = "https://videorouter.sh/api/v1"
H = {"Authorization": f"Bearer {os.environ['VIDEOROUTER_KEY']}",
"Content-Type": "application/json"}
MODELS = ["minimax/h3", "minimax/h3-max"]
PARAMS = {"duration_secs": 5, "resolution": "768p", "aspect_ratio": "16:9"}
REPEATS = 2
MAX_PARALLEL = 3 # stay well under your key's rate limit
def run_one(task):
prompt_id, category, prompt, model, rep = task
base = dict(prompt_id=prompt_id, category=category, model=model, rep=rep)
t0 = time.time()
r = requests.post(f"{API}/videos", headers=H, timeout=30,
json={"model": model, "prompt": prompt, **PARAMS})
if not r.ok:
try:
msg = r.json().get("error", {}).get("message", "")
except ValueError:
msg = r.text[:200]
return dict(base, job_id="", status=f"http_{r.status_code}",
seconds="", url="", error=msg)
job = r.json()
while job["status"] not in ("completed", "failed"):
if time.time() - t0 > 1500:
break
time.sleep(5)
job = requests.get(f"{API}/videos/{job['id']}", headers=H, timeout=30).json()
return dict(base, job_id=job["id"], status=job["status"],
seconds=round(time.time() - t0, 1),
url=(job.get("data") or [{}])[0].get("url", ""),
error=str(job.get("error") or ""))
prompts = json.load(open("prompts.json"))
tasks = [(p["id"], p["category"], p["prompt"], m, rep)
for p in prompts for m in MODELS for rep in range(REPEATS)]
with ThreadPoolExecutor(MAX_PARALLEL) as ex:
rows = list(ex.map(run_one, tasks))
with open("ab_results.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys())
w.writeheader(); w.writerows(rows)
Latency here is wall-clock time from submission to a terminal status, which includes queueing. It is a property of that moment and that host, not a stable number for either model, so repeat at different times of day before reading anything into it.
Recording cost without hard-coding prices
Do not type rates into the script. Compute expected cost from the live table for the tier and host you pinned, multiplied by the requested duration, and reconcile against your balance or usage records afterwards. Billing is once at creation from the requested duration, polling is free, and jobs that fail upstream are not billed, so failed rows should not count toward spend. The live comparison:
| Model | Cheapest host | Priciest host | Cheapest is | Hosts |
|---|---|---|---|---|
| minimax/h3 (768p) | MachGen $0.04 / second | WaveSpeedAI-resell $0.1 / second | 60% lower | 14 |
| minimax/h3-max (480p) | SandBase $0.01 / second | MiniMax $0.05 / second | 80% lower | 5 |
| minimax/h3-unrestricted (768p) | SandBase $0.064 / second | TOAPIS $0.08 / second | 20% lower | 2 |
Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.
Add an expected_cost column that you fill from that table or from your own price file, and treat it as an estimate until checked against actual usage.
Blind review
Write a second script that downloads each url, renames the file to a random token, and stores the token-to-job-id mapping privately. Reviewers see only the tokens. Present the two clips for one prompt side by side, in random left/right order.
import csv, json, os, uuid, requests
os.makedirs("review", exist_ok=True)
rows = [r for r in csv.DictReader(open("ab_results.csv")) if r["status"] == "completed"]
key = {}
for r in rows:
token = uuid.uuid4().hex[:8]
open(f"review/{token}.mp4", "wb").write(requests.get(r["url"]).content)
key[token] = r["job_id"]
json.dump(key, open("blind_key.json", "w")) # keep this away from reviewers
Copy outputs to your own storage promptly rather than treating result URLs as a permanent archive.
A review rubric
Ask each reviewer to score each clip from 1 to 5 on the following, then choose a preferred clip per pair:
| Dimension | Question |
|---|---|
| Prompt adherence | Does the clip contain what the prompt asked for? |
| Motion | Is movement coherent, free of warping or jitter? |
| Subject stability | Do faces, hands and objects keep their shape across frames? |
| Detail | Is fine detail clean at the size you will deliver? |
| Usability | Would you ship it, fix it in post, or regenerate? |
Aggregate by category, not only overall. A pattern such as "no difference except on people" is more actionable than a single winner, because you can route by category in production. Compare preference counts to the price difference on your pinned host before deciding.
What the results can and cannot tell you
- They describe your prompts at one resolution on one host at one time. Rerun after a model update or a host change.
- Failures are data. Count failed jobs and error messages per model; a variant that fails more often costs you retries.
- If the difference is small, the cheaper variant plus a retry budget is a legitimate answer.
Route by what you learned and keep the harness for the next model release. See the model pages for hosts and tiers, the pricing page for the live comparison, and create a key to run it.
Frequently asked questions
How do I compare minimax/h3 and minimax/h3-max fairly?
Send identical prompts and parameters to both, pin the same host so you compare models rather than hosts, repeat important prompts, and have reviewers rank clips blind.
How many prompts do I need?
Ten to twenty representative prompts is a practical start, including your hardest case, with two or three repeats of the prompts that matter most because generation is not deterministic.
Does the harness tell me which model is better?
No. It records job ids, status, latency and outputs so that human reviewers can compare. The conclusion applies to your prompts, one host and one point in time.
Are failed test jobs billed?
Per VideoRouter's documentation, jobs that fail upstream are not billed. Successful jobs are billed once at creation from the requested duration, and polling is free.
Keep reading
- MiniMax H3 Providers Compared — How to Pick a Host
- MiniMax H3 Image-to-Video API: start_image_url Guide
- MiniMax H3 Python Tutorial: Create, Poll, Download, Pin
- MiniMax H3 Resolution Tiers: How Per-Tier Pricing Works
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →