Back to the experimentSupporting notebook

MiniMax H3 — 20-step performance experiments

Source: comfyui/minimaxh3/README.md · revision 6550ead3945b

Actual ComfyUI generation measurements from a three-GPU workstation, preserved with enough provenance to distinguish a genuinely fast result from a mismatched resolution, step count, or quantization profile.

Practical winner: H3 INT8 + Qwen INT8 completed a 960×544, 124-frame, 20-step generation in 354.580 seconds (5m 54.580s). For quality-oriented REF2VA work, the HQ INT8 profile took 399.403 seconds (6m 39.403s)—only 44.823 seconds more.

Results at a glance

Every row produced a 5.167-second H.264/AAC file with 124 frames at 24 fps. All 11 referenced outputs were re-hashed, probed with ffprobe, and fully decoded with ffmpeg -xerror on 2026-08-27.

960×544 cohort

Rank H3 diffusion Text encoder Mode Time Relative note
1 W4A8 pruned NVFP4 REF2VA 5:19.045 Absolute observed winner
2 W4A8 pruned INT8 FL2VA/multishot 5:21.739 Fastest profile with Qwen INT8
3 W4A8 pruned NVFP4 REF2VA 5:31.366 Alternate placement
4 INT8 pruned INT8 T2VA/FL2VA 5:54.580 Recommended when H3 itself must be INT8
5 INT4Q pruned INT8 FL2VA/multishot 6:20.578 Quantized comparison
6 HQ INT8 INT8 REF2VA 6:39.403 Recommended reference-fidelity profile
7 BF16 pruned Q4_K_M REF2VA 8:21.095 BF16 H3 reference point

1344×768 full-resolution cohort

Rank H3 diffusion Text encoder Mode Time Relative note
1 W4A8 pruned W4A8 REF2VA 16:10.306 Fastest observed at full resolution
2 Full INT8, non-pruned INT8 T2VA 16:30.640 Recommended strict Full INT8 path
3 W4A8 pruned INT8 REF2VA 16:57.780 Full-resolution Qwen INT8 comparison
4 Full INT8, non-pruned INT8 REF2VA 17:22.150 Strict Full INT8 reference-to-video

Which profile should I use?

“Full quality” has two independent meanings here. BF16 describes weight precision; 1344×768 describes output resolution. Neither is a perceptual score. These results measure runtime and technical validity, not aesthetic preference, identity fidelity, or prompt adherence.

Repository contents

comfyui/minimaxh3/
├── README.md
├── data/
│   ├── h3_20step_benchmarks.csv
│   ├── h3_20step_benchmarks.json
│   ├── README.md
│   └── runtime_sources.json
├── docs/
│   ├── METHODOLOGY.md
│   └── REPLICATION_AND_BLOG_GUIDE.md
└── scripts/
    ├── check_benchmarks.py
    ├── collect_h3_metrics.py
    ├── validate_local_outputs.py
    └── runtime/
        ├── benchmark_minimax_h3_w4a8_chunks.py
        ├── queue_minimax_h3.py
        ├── run_minimax_h3_gpu_resident.sh
        └── verify_minimax_h3_flash_attention.py

The JSON file is the authoritative dataset. CSV is the compact comparison view. Exact model filenames, GPU placement, VAE choices, sampler, scheduler, workflow SHA-256, output SHA-256, prompt IDs, scene IDs, codecs, dimensions, and throughput are retained. Raw prompts and workflow graphs are deliberately excluded.

The files under scripts/runtime/ are byte-for-byte snapshots of the workstation tools relevant to these measurements: the canonical queue-wrapper entry point, GPU-resident ComfyUI launcher, FlashAttention dispatch verifier, and real-layer W4A8 chunk benchmark. The queue wrapper delegates to the canonical Hermes helper on the source workstation; it is preserved for provenance rather than presented as a standalone portable queue implementation. Their source paths and hashes are recorded in data/runtime_sources.json.

Validate the committed data

From this directory:

python3 scripts/check_benchmarks.py
python3 -m py_compile scripts/*.py scripts/runtime/*.py

On the source workstation, rebuild the dataset from the immutable gallery archive and run a full decode of each output:

python3 scripts/collect_h3_metrics.py --full-decode
python3 scripts/validate_local_outputs.py

The collection script reads WORKSTATION/generation-times.json, archived scene manifests, and the saved ComfyUI workflows. Both roots are configurable with --gallery-root and --comfyui-root.

Important comparison boundary

These are real production runs, not one perfectly controlled A/B matrix. Prompts, generation mode, loader placement, and timing source can differ between rows. Rankings answer “what completed fastest in the recorded experiments?” For causal claims about a quantization or placement change, use only a controlled pair with the same prompt, seed, graph, dimensions, frame count, and runtime state. See the methodology for the complete evidence rules.

For the exact reproduction commands, rationale behind the evidence design, controlled-comparison procedure, and a blog-safe way to describe the result, see the replication and blog guide.

Privacy and storage

This package contains no API keys, tokens, .env files, model weights, rendered videos, prompts, or reference images. Output paths and cryptographic hashes are retained so the source workstation can prove that the locally validated media still matches the recorded experiment.