Actual ComfyUI generation measurements from a three-GPU workstation, preserved with enough provenance to distinguish a genuinely fast result from a mismatched resolution, step count, or quantization profile.
Practical winner: H3 INT8 + Qwen INT8 completed a 960×544, 124-frame, 20-step generation in 354.580 seconds (5m 54.580s). For quality-oriented REF2VA work, the HQ INT8 profile took 399.403 seconds (6m 39.403s)—only 44.823 seconds more.
Results at a glance
Every row produced a 5.167-second H.264/AAC file with 124 frames at 24 fps. All 11 referenced
outputs were re-hashed, probed with ffprobe, and fully decoded with ffmpeg -xerror on
2026-08-27.
960×544 cohort
| Rank | H3 diffusion | Text encoder | Mode | Time | Relative note |
|---|---|---|---|---|---|
| 1 | W4A8 pruned | NVFP4 | REF2VA | 5:19.045 | Absolute observed winner |
| 2 | W4A8 pruned | INT8 | FL2VA/multishot | 5:21.739 | Fastest profile with Qwen INT8 |
| 3 | W4A8 pruned | NVFP4 | REF2VA | 5:31.366 | Alternate placement |
| 4 | INT8 pruned | INT8 | T2VA/FL2VA | 5:54.580 | Recommended when H3 itself must be INT8 |
| 5 | INT4Q pruned | INT8 | FL2VA/multishot | 6:20.578 | Quantized comparison |
| 6 | HQ INT8 | INT8 | REF2VA | 6:39.403 | Recommended reference-fidelity profile |
| 7 | BF16 pruned | Q4_K_M | REF2VA | 8:21.095 | BF16 H3 reference point |
1344×768 full-resolution cohort
| Rank | H3 diffusion | Text encoder | Mode | Time | Relative note |
|---|---|---|---|---|---|
| 1 | W4A8 pruned | W4A8 | REF2VA | 16:10.306 | Fastest observed at full resolution |
| 2 | Full INT8, non-pruned | INT8 | T2VA | 16:30.640 | Recommended strict Full INT8 path |
| 3 | W4A8 pruned | INT8 | REF2VA | 16:57.780 | Full-resolution Qwen INT8 comparison |
| 4 | Full INT8, non-pruned | INT8 | REF2VA | 17:22.150 | Strict Full INT8 reference-to-video |
Which profile should I use?
- Fast 20-step INT8: use
h3-int8-qwen-int8-fl2va-960x544. It is the fastest measured run where both the H3 diffusion model and Qwen encoder are INT8. - Reference-image quality: use
h3-hq-int8-qwen-int8-ref2va-960x544. It costs 12.64% more wall time than the speed-oriented INT8 profile and retains the HQ REF2VA model. - Strict full-resolution Full INT8: use
full-int8-t2va-1344x768, or the REF2VA counterpart when a reference image is required. - Absolute minimum time: W4A8 + NVFP4 reached 319.045 seconds, but that is not an INT8 H3 profile and should not be described as one.
“Full quality” has two independent meanings here. BF16 describes weight precision; 1344×768 describes output resolution. Neither is a perceptual score. These results measure runtime and technical validity, not aesthetic preference, identity fidelity, or prompt adherence.
Repository contents
comfyui/minimaxh3/
├── README.md
├── data/
│ ├── h3_20step_benchmarks.csv
│ ├── h3_20step_benchmarks.json
│ ├── README.md
│ └── runtime_sources.json
├── docs/
│ ├── METHODOLOGY.md
│ └── REPLICATION_AND_BLOG_GUIDE.md
└── scripts/
├── check_benchmarks.py
├── collect_h3_metrics.py
├── validate_local_outputs.py
└── runtime/
├── benchmark_minimax_h3_w4a8_chunks.py
├── queue_minimax_h3.py
├── run_minimax_h3_gpu_resident.sh
└── verify_minimax_h3_flash_attention.py
The JSON file is the authoritative dataset. CSV is the compact comparison view. Exact model filenames, GPU placement, VAE choices, sampler, scheduler, workflow SHA-256, output SHA-256, prompt IDs, scene IDs, codecs, dimensions, and throughput are retained. Raw prompts and workflow graphs are deliberately excluded.
The files under scripts/runtime/ are byte-for-byte snapshots of the workstation tools relevant
to these measurements: the canonical queue-wrapper entry point, GPU-resident ComfyUI launcher,
FlashAttention dispatch verifier, and real-layer W4A8 chunk benchmark. The queue wrapper delegates
to the canonical Hermes helper on the source workstation; it is preserved for provenance rather
than presented as a standalone portable queue implementation. Their source paths and hashes are
recorded in data/runtime_sources.json.
Validate the committed data
From this directory:
python3 scripts/check_benchmarks.py
python3 -m py_compile scripts/*.py scripts/runtime/*.py
On the source workstation, rebuild the dataset from the immutable gallery archive and run a full decode of each output:
python3 scripts/collect_h3_metrics.py --full-decode
python3 scripts/validate_local_outputs.py
The collection script reads WORKSTATION/generation-times.json, archived scene
manifests, and the saved ComfyUI workflows. Both roots are configurable with --gallery-root and
--comfyui-root.
Important comparison boundary
These are real production runs, not one perfectly controlled A/B matrix. Prompts, generation mode, loader placement, and timing source can differ between rows. Rankings answer “what completed fastest in the recorded experiments?” For causal claims about a quantization or placement change, use only a controlled pair with the same prompt, seed, graph, dimensions, frame count, and runtime state. See the methodology for the complete evidence rules.
For the exact reproduction commands, rationale behind the evidence design, controlled-comparison procedure, and a blog-safe way to describe the result, see the replication and blog guide.
Privacy and storage
This package contains no API keys, tokens, .env files, model weights, rendered videos, prompts,
or reference images. Output paths and cryptographic hashes are retained so the source workstation
can prove that the locally validated media still matches the recorded experiment.