Objective
Identify the fastest actual MiniMax H3 generations that used 20 sampling steps while keeping resolution, frame count, model precision, and generation mode visible. The dataset is designed to prevent a 960×544 quantized run from being presented as directly equivalent to a 1344×768 Full INT8 or BF16 run.
Source hierarchy
WORKSTATION/generation-times.jsonsupplies the recorded elapsed time, label, output path, prompt ID, and measurement method.- Immutable manifests under
scene-archive/scenes/supply the exact executed ComfyUI graph for archived jobs. - Saved API workflows under
WORKSTATION/adaptedcover verified historical rows that predate durable scene archiving. - The local output beneath
WORKSTATION/outputsupplies codec, duration, dimensions, frame count, file size, and output SHA-256.
The exporter parses only operational graph fields. It does not export prompt text, reference-image names, media bytes, model weights, or credentials.
Inclusion criteria
A row is included only when all of the following are true:
- MiniMax H3 workflow with an exact graph or saved workflow source;
- exactly 20 sampling steps;
- 124 output video frames at 24 fps;
- 5.167-second output with H.264 video and AAC audio;
- width and height belong to the declared 960×544 or 1344×768 cohort;
- elapsed time exists in the gallery timing ledger;
- output exists locally and its SHA-256 can be calculated;
ffprobesucceeds and a fullffmpeg -v error -xerror ... -map 0 -f null -decode passes.
Timing semantics
Most archived rows use ComfyUI execution_start to execution_success timestamps. A few older
verified rows use queue-helper wall time. The measurement field preserves this distinction.
Wall-clock timing may include startup or finalization work that is outside the execution event
window, so small differences between those timing methods are not attributed solely to model
precision.
Cohorts
960x544-124f-20step: 522,240 pixels per frame.1344x768-124f-20step: 1,032,192 pixels per frame, approximately 1.98× the pixels of the lower resolution.
Rows are ranked globally for discoverability and separately within their resolution cohort. The cohort still contains different modes and placements, so it is a ranking of observed runs rather than a universal quantization benchmark.
Precision labels
- BF16: BF16 diffusion weights.
- INT8: INT8 ConvRot diffusion or text-encoder weights.
- W4A8: weight-4/activation-8 mixed quantization.
- INT4Q: the local INT4Q diffusion variant.
- NVFP4: NVIDIA FP4 text-encoder variant.
- Q4_K_M: GGUF Q4_K_M text encoder.
diffusion_pruned is recorded independently. “Full INT8” in the results means a non-pruned INT8
H3 diffusion graph at 1344×768; it does not imply that every auxiliary component uses the same
precision.
Derived metrics
seconds_per_output_second = elapsed_seconds / output_duration_secondsgenerated_frames_per_second = output_frames / elapsed_secondscohort_ranksorts elapsed time within the declared resolution cohort.
Derived values are rounded only after calculation. The raw elapsed and output duration remain in the JSON dataset.
Reproduction and validation
cd comfyui/minimaxh3
python3 scripts/collect_h3_metrics.py --full-decode
python3 scripts/check_benchmarks.py
python3 scripts/validate_local_outputs.py
collect_h3_metrics.py re-reads the live ledger/workflows, recalculates workflow and output
SHA-256 values, probes each MP4, and optionally fully decodes it. check_benchmarks.py is portable
and runs in CI against the committed JSON/CSV. validate_local_outputs.py requires the source
workstation’s media archive and verifies that current files still match the committed hashes.
What this dataset does not prove
- A speed difference is not automatically caused by quantization alone.
- Technical decode success is not a perceptual quality score.
- Prompt adherence, identity preservation, lip sync, and visual artifacts require a separate blinded or visible review.
- A measured run is not a performance SLA; model residency, caches, GPU temperature, and other workloads can change elapsed time.