Journal

The real cost of a five-second video

At 960×544, dual INT8 completed a 124-frame, 20-step run in 354.580 seconds.

Vlad / experimentos.
The boundary

Runs vary in mode, quantization and placement. Separate resolution cohorts are observed production runs, not one controlled A/B.

Measurements

Recorded result

A five-second video at 960×544

Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed

A five-second video at 960×544W4A8 / NVFP4 · REF2VA: 319.045 elapsed seconds; W4A8 / INT8 · multishot: 321.739 elapsed seconds; W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed seconds; INT4Q / INT8 · multishot: 380.578 elapsed seconds; INT8 / INT8 · REF2VA · HQ: 399.403 elapsed seconds; BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.W4A8 / NVFP4 · REF2VA319.045W4A8 / NVFP4 · REF2VA: 319.045 elapsed secondsW4A8 / INT8 · multishot321.739W4A8 / INT8 · multishot: 321.739 elapsed secondsW4A8 / NVFP4 · REF2VA · alt. placement331.366W4A8 / NVFP4 · REF2VA · alt. placement: 331.366 elapsed secondsINT8 / INT8 · T2VA/FL2VA354.580INT8 / INT8 · T2VA/FL2VA: 354.580 elapsed secondsINT4Q / INT8 · multishot380.578INT4Q / INT8 · multishot: 380.578 elapsed secondsINT8 / INT8 · REF2VA · HQ399.403INT8 / INT8 · REF2VA · HQ: 399.403 elapsed secondsBF16 / Q4_K_M · REF2VA501.095BF16 / Q4_K_M · REF2VA: 501.095 elapsed seconds0elapsed seconds
  1. W4A8 / NVFP4 · REF2VA319.045
  2. W4A8 / INT8 · multishot321.739
  3. W4A8 / NVFP4 · REF2VA · alt. placement331.366
  4. INT8 / INT8 · T2VA/FL2VA354.580
  5. INT4Q / INT8 · multishot380.578
  6. INT8 / INT8 · REF2VA · HQ399.403
  7. BF16 / Q4_K_M · REF2VA501.095

elapsed seconds

Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.

View data & source
A five-second video at 960×544 · elapsed seconds
ConfigurationValue
W4A8 / NVFP4 · REF2VA319.045
W4A8 / INT8 · multishot321.739
W4A8 / NVFP4 · REF2VA · alt. placement331.366
INT8 / INT8 · T2VA/FL2VA354.580
INT4Q / INT8 · multishot380.578
INT8 / INT8 · REF2VA · HQ399.403
BF16 / Q4_K_M · REF2VA501.095

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

A five-second video at 1344×768

Three RTX 3090s · 124 frames · 24 fps · 20 steps · full-decode checks passed

A five-second video at 1344×768W4A8 / W4A8 · REF2VA: 970.306 elapsed seconds; INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed seconds; W4A8 / INT8 · REF2VA: 1,017.780 elapsed seconds; INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds. Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.W4A8 / W4A8 · REF2VA970.306W4A8 / W4A8 · REF2VA: 970.306 elapsed secondsINT8 / INT8 · T2VA/FL2VA990.640INT8 / INT8 · T2VA/FL2VA: 990.640 elapsed secondsW4A8 / INT8 · REF2VA1,017.780W4A8 / INT8 · REF2VA: 1,017.780 elapsed secondsINT8 / INT8 · REF2VA1,042.150INT8 / INT8 · REF2VA: 1,042.150 elapsed seconds0elapsed seconds
  1. W4A8 / W4A8 · REF2VA970.306
  2. INT8 / INT8 · T2VA/FL2VA990.640
  3. W4A8 / INT8 · REF2VA1,017.780
  4. INT8 / INT8 · REF2VA1,042.150

elapsed seconds

Observed production runs: mode, placement and measurement method vary. Read each profile before choosing.

View data & source
A five-second video at 1344×768 · elapsed seconds
ConfigurationValue
W4A8 / W4A8 · REF2VA970.306
INT8 / INT8 · T2VA/FL2VA990.640
W4A8 / INT8 · REF2VA1,017.780
INT8 / INT8 · REF2VA1,042.150

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

Five seconds of generated video can take many minutes of compute. The H3 archive groups real 20-step runs by resolution and keeps quantization, placement, mode and complete decode checks attached to each timing.

From the original notebook

Actual ComfyUI generation measurements from a three-GPU workstation, preserved with enough provenance to distinguish a genuinely fast result from a mismatched resolution, step count, or quantization profile.

Practical winner: H3 INT8 + Qwen INT8 completed a 960×544, 124-frame, 20-step generation in 354.580 seconds (5m 54.580s). For quality-oriented REF2VA work, the HQ INT8 profile took 399.403 seconds (6m 39.403s)—only 44.823 seconds more.

Results at a glance

Every row produced a 5.167-second H.264/AAC file with 124 frames at 24 fps. All 11 referenced outputs were re-hashed, probed with ffprobe, and fully decoded with ffmpeg -xerror on 2026-08-27.

960×544 cohort

Rank H3 diffusion Text encoder Mode Time Relative note
1 W4A8 pruned NVFP4 REF2VA 5:19.045 Absolute observed winner
2 W4A8 pruned INT8 FL2VA/multishot 5:21.739 Fastest profile with Qwen INT8
3 W4A8 pruned NVFP4 REF2VA 5:31.366 Alternate placement
4 INT8 pruned INT8 T2VA/FL2VA 5:54.580 Recommended when H3 itself must be INT8
5 INT4Q pruned INT8 FL2VA/multishot 6:20.578 Quantized comparison
6 HQ INT8 INT8 REF2VA 6:39.403 Recommended reference-fidelity profile
7 BF16 pruned Q4_K_M REF2VA 8:21.095 BF16 H3 reference point

1344×768 full-resolution cohort

Rank H3 diffusion Text encoder Mode Time Relative note
1 W4A8 pruned W4A8 REF2VA 16:10.306 Fastest observed at full resolution
2 Full INT8, non-pruned INT8 T2VA 16:30.640 Recommended strict Full INT8 path
3 W4A8 pruned INT8 REF2VA 16:57.780 Full-resolution Qwen INT8 comparison
4 Full INT8, non-pruned INT8 REF2VA 17:22.150 Strict Full INT8 reference-to-video

Which profile should I use?

  • Fast 20-step INT8: use h3-int8-qwen-int8-fl2va-960x544. It is the fastest measured run where both the H3 diffusion model and Qwen encoder are INT8.
  • Reference-image quality: use h3-hq-int8-qwen-int8-ref2va-960x544. It costs 12.64% more wall time than the speed-oriented INT8 profile and retains the HQ REF2VA model.
  • Strict full-resolution Full INT8: use full-int8-t2va-1344x768, or the REF2VA counterpart when a reference image is required.
  • Absolute minimum time: W4A8 + NVFP4 reached 319.045 seconds, but that is not an INT8 H3 profile and should not be described as one.

“Full quality” has two independent meanings here. BF16 describes weight precision; 1344×768 describes output resolution. Neither is a perceptual score. These results measure runtime and technical validity, not aesthetic preference, identity fidelity, or prompt adherence.

Repository contents

comfyui/minimaxh3/
├── README.md
├── data/
│   ├── h3_20step_benchmarks.csv
│   ├── h3_20step_benchmarks.json
│   ├── README.md
│   └── runtime_sources.json
├── docs/
│   ├── METHODOLOGY.md
│   └── REPLICATION_AND_BLOG_GUIDE.md
└── scripts/
    ├── check_benchmarks.py
    ├── collect_h3_metrics.py
    ├── validate_local_outputs.py
    └── runtime/
        ├── benchmark_minimax_h3_w4a8_chunks.py
        ├── queue_minimax_h3.py
        ├── run_minimax_h3_gpu_resident.sh
        └── verify_minimax_h3_flash_attention.py

The JSON file is the authoritative dataset. CSV is the compact comparison view. Exact model filenames, GPU placement, VAE choices, sampler, scheduler, workflow SHA-256, output SHA-256, prompt IDs, scene IDs, codecs, dimensions, and throughput are retained. Raw prompts and workflow graphs are deliberately excluded.

The files under scripts/runtime/ are byte-for-byte snapshots of the workstation tools relevant to these measurements: the canonical queue-wrapper entry point, GPU-resident ComfyUI launcher, FlashAttention dispatch verifier, and real-layer W4A8 chunk benchmark. The queue wrapper delegates to the canonical Hermes helper on the source workstation; it is preserved for provenance rather than presented as a standalone portable queue implementation. Their source paths and hashes are recorded in data/runtime_sources.json.

Validate the committed data

From this directory:

python3 scripts/check_benchmarks.py
python3 -m py_compile scripts/*.py scripts/runtime/*.py

On the source workstation, rebuild the dataset from the immutable gallery archive and run a full decode of each output:

python3 scripts/collect_h3_metrics.py --full-decode
python3 scripts/validate_local_outputs.py

The collection script reads WORKSTATION/generation-times.json, archived scene manifests, and the saved ComfyUI workflows. Both roots are configurable with --gallery-root and --comfyui-root.

Important comparison boundary

These are real production runs, not one perfectly controlled A/B matrix. Prompts, generation mode, loader placement, and timing source can differ between rows. Rankings answer “what completed fastest in the recorded experiments?” For causal claims about a quantization or placement change, use only a controlled pair with the same prompt, seed, graph, dimensions, frame count, and runtime state. See the methodology for the complete evidence rules.

For the exact reproduction commands, rationale behind the evidence design, controlled-comparison procedure, and a blog-safe way to describe the result, see the replication and blog guide.

Privacy and storage

This package contains no API keys, tokens, .env files, model weights, rendered videos, prompts, or reference images. Output paths and cryptographic hashes are retained so the source workstation can prove that the locally validated media still matches the recorded experiment.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (4)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS