Back to the experimentSupporting notebook

Replication guide and blog notes

Source: comfyui/minimaxh3/docs/REPLICATION_AND_BLOG_GUIDE.md · revision 6550ead3945b

This document explains the engineering choices behind the MiniMax H3 benchmark, how to revalidate the published evidence, and how to describe the result accurately in a future post. It deliberately separates facts measured on the source workstation from conclusions that would need a controlled A/B experiment.

The question this experiment answers

The practical question was: of the real completed 20-step MiniMax H3 generations, which profile finished fastest, and which INT8 option is the best speed/quality trade-off?

The result is not a claim that one quantization is universally better. It is an evidence-backed ranking of recorded production runs, with enough context to avoid comparing a small, pruned, reference-to-video run to a full-resolution, non-pruned text-to-video run as if they were the same job.

Decisions and the reasons for them

Decision Why it matters
Keep the generation mode, resolution, frame count, pruning state, and both model precisions in every row. A shorter job or a different graph can appear faster for reasons unrelated to the precision being compared.
Use the gallery timing ledger plus immutable archived graphs or saved workflows. Timing alone is not enough: the graph establishes what actually ran.
Require a local output hash, ffprobe, and a full ffmpeg -xerror decode. A file existing or playing in one client is weaker evidence than a complete decode of its streams.
Publish hashes and operational metadata, but not prompts, media, reference images, weights, or credentials. This preserves provenance and lets the source workstation revalidate outputs without exposing private creative material or large artifacts.
Rank both globally and by resolution cohort. The global winner is useful, while cohort ranking makes the fairer same-output-size comparison easy to see.
Keep the runtime scripts and their source hashes. A benchmark result is more useful when the launcher, queue path, and dispatch checks that produced it can be inspected later.

Snapshot of the measured result

All rows are 20-step, 124-frame, 24-fps outputs (5.167 seconds) and passed the recorded media validation. The detailed table is in the experiment README.

Question Observed answer Interpretation boundary
Absolute fastest recorded 960×544 result W4A8 H3 + NVFP4 Qwen, REF2VA: 319.045 s It is not an INT8-H3 profile.
Fastest recorded result with Qwen INT8 W4A8 H3 + INT8 Qwen, FL2VA: 321.739 s The diffusion model is W4A8, not INT8.
Fastest recorded result with both H3 and Qwen INT8 Pruned H3 INT8 + Qwen INT8, T2VA/FL2VA: 354.580 s Best answer when “INT8” means both main models must be INT8.
Quality-oriented INT8 REF2VA profile HQ INT8 H3 + Qwen INT8: 399.403 s A reference-image workflow; technical validity is not a perceptual quality score.
Strict Full INT8 at 1344×768 Non-pruned H3 INT8 + Qwen INT8, T2VA: 990.640 s A distinct, nearly 2×-pixel cohort.

The useful operational recommendation is therefore:

Revalidate the committed package

These commands need only Python 3 and the files in this repository. They verify that the machine-readable data, CSV projection, and recorded runtime-script hashes are internally consistent; they do not need the original media archive.

git clone git@github.com:groxaxo/experimentos.git
cd experimentos/comfyui/minimaxh3
python3 scripts/check_benchmarks.py
python3 -m py_compile scripts/*.py scripts/runtime/*.py
bash -n scripts/runtime/run_minimax_h3_gpu_resident.sh

Expected result: check_benchmarks.py reports that all 11 curated runs are valid.

Rebuild and revalidate on the source workstation

The collector defaults are intentionally source-workstation specific:

With that archive present, regenerate the dataset and prove the media bytes are still the ones recorded here:

cd WORKSTATION/minimaxh3
python3 scripts/collect_h3_metrics.py --full-decode
python3 scripts/check_benchmarks.py
python3 scripts/validate_local_outputs.py

--full-decode invokes ffmpeg -v error -xerror -map 0 -f null - for every selected output. validate_local_outputs.py additionally compares every current output SHA-256 against the committed dataset. Both commands should be run after copying or restoring media, because a successful ffprobe alone does not establish byte identity or full decodability.

To use a different checkout or a restored archive, point the collector at it explicitly:

python3 scripts/collect_h3_metrics.py \
  --gallery-root /path/to/media-gallery \
  --comfyui-root /path/to/ComfyUI \
  --full-decode

The collector is intentionally conservative: it exports only the curated profiles described in the code and refuses incomplete evidence instead of silently filling it with guesses.

Run a new controlled comparison

The eleven historical runs are excellent operational evidence, but they are not a causal quantization study. For a claim such as “INT8 is faster than W4A8 by X%,” make a new controlled matrix.

  1. Start from one known workflow and change only the variable being tested. Keep prompt, seed, reference input, sampler, scheduler, steps, dimensions, frames, FPS, VAE, and generation mode identical.

  2. Record the exact model filenames, graph SHA-256, ComfyUI revision, launcher arguments, GPU placement, and whether the models were warm/resident before each run.

  3. Drain the queue and avoid unrelated GPU work. Capture nvidia-smi before and after each run; temperature, clocks, cache state, and contention can move wall time materially.

  4. Run multiple interleaved repeats rather than all of profile A followed by all of profile B. Report the raw samples plus median and spread, not only the best run.

  5. Use the preserved launcher and dispatch tools as the starting point, then record any changes:

    scripts/runtime/run_minimax_h3_gpu_resident.sh --help
    python3 scripts/runtime/queue_minimax_h3.py --help
    python3 scripts/runtime/verify_minimax_h3_flash_attention.py --help
  6. For every output, calculate SHA-256, run ffprobe, and run the full ffmpeg -xerror decode before including its elapsed time. Add the new evidence to a separately named dataset rather than rewriting this historical snapshot.

The preserved queue wrapper delegates to the source workstation’s canonical helper. It is a provenance snapshot, not a promise that it can queue a job unchanged on another host. Read its --help, verify the model inventory, and perform a low-cost smoke before any expensive job.

A defensible blog narrative

A concise, accurate framing is:

On a three-GPU local ComfyUI workstation, I revalidated eleven real MiniMax H3 20-step outputs with workflow provenance, SHA-256 hashes, stream probing, and full FFmpeg decoding. The fastest observed 960×544 run was W4A8 H3 plus NVFP4 Qwen at 319 seconds. When both H3 and Qwen had to be INT8, the best recorded run was 355 seconds; the HQ INT8 reference-to-video profile was 399 seconds.

Useful points to explain:

Avoid saying that INT8 is “the fastest” without qualification, that the profiles have identical visual quality, or that the timings are a universal service-level guarantee. Say “observed on this workstation” and include the resolution, frame count, steps, mode, and precision profile beside every headline number.

For a public post, publish a sanitized version of the result table and this methodology, but omit absolute private paths, prompt text, reference media, model files, tokens, and output media unless you have a deliberate release policy for each. The hashes, dimensions, timing semantics, and workflow hashes are enough to communicate reproducibility without exposing the creative inputs.

Validation and CI note

The package is validated locally with the commands above. An optional GitHub Actions workflow was intentionally not retained because the private account’s hosted runner was unavailable due to a billing/spending gate. That is a hosting limitation, not a substitute for validation; every published revision should run the portable checker and the source-workstation media validation before it is described as refreshed.