Back to the experimentSupporting notebook

Qwen-Image-2.1: TE AutoRound-AWQ on GPU + all-resident single-3090 stack

Source: qwen21-te-awq/README.md · revision 6550ead3945b

Date: 2026-09-21 · Workstation: 3× RTX 3090 24 GB, torch 2.14.0+cu130, diffusers 0.41.0.dev0, transformers 5.17, autoround 0.15.1, autoawq 0.2.9, torchao (2026-09 build)

Default configuration — quality-first, 2026-09-24

The committed default keeps the existing AWQ text encoder and FP8 DiT weight precision, with Triton AWQ kernels, regional DiT compilation, native attention, and an eager BF16 VAE using 512-pixel tiles. It keeps 1024×1024, 40 steps and true CFG 4.0. No dynamic INT8 activation quantization or VAE compilation is enabled implicitly. This is the previously validated 93.686-second warm configuration, not the 46-second experimental configuration below.

## Check availability first; use a free GPU only.
nvidia-smi
PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 \
  WORKSTATION/python scripts/gpu_resident_test.py \
  --output-dir /tmp/qwen21-quality-default

“Quality-first” means no additional precision reduction relative to the existing AWQ/FP8 stack, not lossless equivalence to the original BF16 model. Compilation, kernel selection and VAE tiling can still change pixels. The two-prompt visual checks and conditioning comparisons below do not prove zero quality loss across all prompts. In particular, we cannot promise the 46-second INT8 path without a quality tradeoff, so it remains an explicit opt-in.

Check the CLI defaults without loading a model or using a GPU:

python3 -m unittest discover -s scripts -p 'test_benchmark_defaults.py' -v

The initial 2026-09-24 release check found an incomplete local Hugging Face snapshot. The two required DiT shards were subsequently restored from the pinned revision, allowing the expanded offline generation suite below to complete. The two-prompt GPU measurements immediately below remain the 2026-09-23 runs; their FP8 defaults and quality policy were preserved. Committed generation logs retain their output with trailing terminal padding removed for clean diffs; numeric results are unchanged.

Whole-pipeline follow-up — 2026-09-23

An opt-in W8A8 DiT path cuts the previous optimized warm time in half on the same single RTX 3090: 46.024 seconds mean, including compiled VAE decoding, with 17.3 GiB peak allocated. The AWQ text encoder stays on its verified Triton backend; the DiT’s 224 block linear layers now use INT8 weights and dynamic INT8 activations, with regional compilation. Eight boundary/conditioning linear layers and the tiled VAE remain BF16. This is a different quantization scheme, not a claim that the entire pipeline has become four-bit AWQ.

Configuration Seed 42 Seed 7 Mean Peak allocated
Original eager FP8/ATen 113.876 s 114.043 s 113.960 s 18.4 GiB
Previous compiled FP8/Triton 93.557 s 93.814 s 93.686 s 17.5 GiB
Compiled INT8/Triton, native attention 46.635 s 46.394 s 46.515 s 17.7 GiB
INT8 + compiled VAE, native attention 45.927 s 46.121 s 46.024 s 17.3 GiB

All rows use the same two prompts/seeds, 1024×1024, 40 steps, true CFG 4.0. Each timed image still performs 80 DiT forwards. Loading, compilation and the two-step prompt-specific warm-up are excluded; all reported measurements had zero newly compiled graphs. The INT8 measurements used physical GPU 0 only, with its existing 245 W power cap. Other GPU workloads were left untouched. These are two-sample measurements, not a broad performance or quality study.

Run the faster, numerically different path:

PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 \
  WORKSTATION/python scripts/gpu_resident_test.py \
  --dit-weights int8-dynamic --compile-vae --output-dir /tmp/qwen21-int8

Use only a GPU that is free at launch. FP8 remains the default: INT8 changes fine image details and is intentionally opt-in. Both full-step samples retain their requested content (including the exact neon text), with no obvious visual collapse or tile seams, but this does not establish lossless or universally equivalent quality. Against the previous compiled FP8 samples, final RGB MAE is 11.73 / 8.56 out of 255, and PSNR is 18.22 / 23.24 dB. These metrics measure image difference, not aesthetic quality. The expanded text-to-image suite below broadens the visual check. Image-editing validation has not been performed; the slim text-only encoder does not support image edits. Omit --compile-vae for the simpler 46.515-second variant. Neither variant is pixel-identical to the previous FP8 configuration or to each other.

Profiling the previous FP8 path identified its large BF16 matrix multiplications as the main cost: two CUTLASS kernel families alone accounted for 65.85% of all CUDA kernel time in a two-step diagnostic run. FP8 was saving weight memory but still doing BF16 matrix multiplication. The INT8 path changes that arithmetic; the loader verifies that every DiT block linear has an Int8Tensor weight with activation quantization enabled. TorchAO’s INT8 configuration uses per-token activation and per-channel weight quantization. The actual run used installed TorchAO 0.18.0 configuration version 2; website signatures may describe an older release.

The final kernel trace confirms 896 INT8 CUTLASS GEMMs for the two-step diagnostic (224 block linears × four CFG forwards), plus Triton fused activation quantization, normalization, activation and VAE surrounding operations. The encoder executes awq_gemm_kernel. Attention remains native FlashAttention and VAE convolutions remain cuDNN: this is a measured mixed-kernel optimization, not an all-Triton rewrite. See the final kernel summary.

Evidence: final full-step results, final configuration, final image differences, final images, INT8 with eager VAE results, and FP8 kernel profile summary. Raw generation logs are beside the results. No checkpoint or installed package was modified; quantization happens in memory during loading.

Kernel-selection probes

Six-step, single-prompt probes (not substitutes for the full-step results):

Probe Warm total DiT VAE decode
FP8 + maximum matrix-multiply autotuning 15.048 s 13.986 s 0.914 s
INT8 + native attention 7.907 s 6.844 s 0.913 s
INT8 + Triton Flex attention 7.913 s 6.866 s 0.907 s
INT8 + compiled VAE, native attention 7.530 s 6.775 s 0.626 s

Native attention is already using PyTorch FlashAttention CUDA kernels. Flex was not measurably faster in this probe, so native remains selected. Maximum GEMM autotuning also did not improve the existing FP8 path meaningfully. These options remain available for diagnostics as --attention flex and --compile-mode max-autotune-no-cudagraphs; neither is the recommended default.

--compile-vae compiles the decoder’s forward with timing hooks kept outside the graph. TorchAO enables coordinate-descent tuning globally; allowing that inside the VAE left a probe warming when interrupted after 6½ minutes wall time. The implemented VAE path disables that tuning locally. Its subsequent warm-up took 50.95 seconds with partially populated caches, not a clean-cache cold-start measurement. First-time compilation can still take minutes. The VAE stays BF16 and tiled; compilation is not VAE quantization.

To inspect actual kernels, add --profile --steps 2 --jobs 1, then summarize profile_s42.json with scripts/summarize_kernel_profile.py TRACE --output JSON. Profiler runs are marked profiled-warm and must not be used as ordinary latency results. The summarizer counts CUDA kernel events only, excluding overlapping profiler annotations. Probe JSON/log files are in data/full-triton-20260923/.

Speed follow-up — 2026-09-23

The benchmark now defaults to Triton AWQ text encoding + regionally compiled FP8 DiT + tiled BF16 VAE, all resident on one GPU. Two warmed 1024×1024, 40-step, true-CFG 4.0 images took 93.557 / 93.814 seconds, with 17.5 GiB peak allocated. The earlier explicitly warmed run took 113.2 / 113.3 seconds. No steps or guidance passes were removed: the measured run still executes 80 DiT forwards per image.

Fresh paired measurements on physical GPU 0 (RTX 3090, configured 245 W power limit):

Seed Baseline warm seconds Optimized warm seconds
42 113.876 93.557
7 114.043 93.814
Mean 113.960 93.686

That is 17.8% less latency / 1.22× throughput. Peak allocated VRAM is 18.4 GiB for the baseline after loader cleanup, versus 17.5 GiB optimized. These are two samples rather than a statistically powered throughput study.

Run on a GPU that is free at launch (GPU 0 was used for this validation):

PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 \
  WORKSTATION/python scripts/gpu_resident_test.py \
  --output-dir /tmp/qwen21-optimized

## Same inputs through the previous execution path:
PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 \
  WORKSTATION/python scripts/gpu_resident_test.py \
  --compile none --awq-backend aten --no-vae-tiling \
  --output-dir /tmp/qwen21-baseline

## Recheck actual pipeline conditioning for both AWQ kernels:
PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 \
  WORKSTATION/python scripts/benchmark_awq_backends.py \
  --output /tmp/qwen21-awq-backends.json

PYTHONNOUSERSITE=1 WORKSTATION/python \
  scripts/compare_benchmark_images.py /tmp/qwen21-baseline /tmp/qwen21-optimized \
  --output /tmp/qwen21-image-comparison.json

Outputs include both PNGs and results.json; the final harness also saves config.json with package versions and GPU metadata. Set TORCHINDUCTOR_COMPILE_THREADS to override the default four compile workers. Two steps warm the tested prompt shapes; first-time compilation can take substantially longer than a cached launch.

The images are not pixel-identical after compilation/kernel changes and tiled decoding. Visual inspection of both optimized images confirms the neon text and cabin scene, without obvious tile seams. This is a two-prompt regression check, not a broad quality evaluation. The fresh baseline images exactly match the stored samples. Optimized RGB mean absolute differences are 2.06 / 2.91 on the 0–255 scale (PSNR 30.01 / 30.28 dB); alpha mean differences are 0.019 / 0.066. The older prompt-fidelity caveat below is historical: the stored and newly generated seed-42/seed-7 samples do show the requested sign and cabin.

Rejected configurations: compilation with untiled VAE decoding ran out of memory; BF16 DiT weights with 256-pixel VAE tiles completed a six-step test but peaked at 23.0 GiB for only a small extra speed gain. The latter remains an explicitly experimental --dit-weights bf16 option. FP8 here is still weight storage with BF16 matrix multiplication; this change does not enable native FP8 on the 3090.

Evidence: baseline, optimized, AWQ backend comparison, pixel comparison, and optimized images. Raw generation logs are stored beside the JSON results. The final two-step smoke results validate the default options and zero new graphs during timed generation; they are not the 40-step performance measurements above.

Question answered

Can the full Qwen-Image-2.1 stack (Qwen3-VL-8B text encoder + 7B DiT + VAE) run all-resident on one 24 GB RTX 3090, with the text encoder quantized to int4 AWQ running on the GPU (not CPU), at validated quality — and what breaks along the way?

Key measured result

Config s/image Peak VRAM Verdict
bf16 TE + DiT bf16, enable_model_cpu_offload 68.3 (offloaded) baseline, slow
fp8 DiT + offload (bf16 TE) 71.0 (offloaded) baseline, slow
int8 TE + fp8 DiT resident ~110 22.6 flaky, OOM cliff
w4 ARK TE on CPU + fp8 DiT 124–153 13.9 stable, CPU encode tax
w4 AWQ TE slim on GPU + fp8 DiT resident 110.7 / 111.4 19.6 stable, margin >3 GiB

Sample evidence (same prompt/seed pairs across quants): samples/gpu_res_s42.png, samples/gpu_res_s7.png (final stack) vs samples/bf16te_fp8_offload_s42.png (bf16 baseline) and samples/combo_s42.png (int8 TE). Images delivered to Telegram (message ids 12663/12664/12675/12676).

The journey (what broke and why)

  1. 2-GPU device_map="balanced" → flat gray noise (see ../qwen21_2gpu/BROKEN.md). Cross-GPU TE shards corrupt conditioning for this pipeline. Single-GPU only.
  2. TE w4 (ARK packing auto_round:auto_awq) cannot move to CUDA: ARK kernels are cpu/xpu-native — blob must reside on cpu or xpu, got cuda. This is what forced the intermediate CPU-shim design (scripts/w4te_cpu_shim.py): TE on CPU, embeddings bridged to CUDA bf16.
  3. Resident int8 TE + fp8 DiT is memory-marginal: 20.8–23.5 GiB peaks on a 23.56 GiB card; flaky OOM on the first DiT forward. Root cause was the memory cliff, not the (already-fixed) upstream KV-cache .contiguous() bug — this diffusers build has cache_write_slice].clone().
  4. Swapping pipe.text_encoder breaks pipe.device/pipe._execution_device in this diffusers build (ConfigMixin __getattr__). Fix: give the shim .device/.dtype properties — then pipe.device resolves cuda naturally, no class-level monkey patch needed.
  5. Qwen3VLCausalLMOutputWithPast has no .to() and dtype mismatch (fp16 TE output vs bf16 DiT, Half != BFloat16 at img_in): the shim calls te.model(...) directly (skips lm_head), returns only hidden_states[-1] cast to bf16. The pipeline’s final-RMSNorm bypass hook still fires correctly through the shim’s .model property.
  6. The raw text_encoder tokenizer emits EMPTY input_ids for plain text (even "hello" → [1, 0]). Tokenization must go through the processor with chat template — this silently broke a validation script with cannot reshape tensor of 0 elements. scripts/diag_tok.py reproduces it.
  7. Meta-tensor order matters on torch 2.14: .to("meta") must come after .to("cuda") placement (moving meta tensors with .to() raises NotImplementedError).

Requant to AWQ (GPU-usable packing)

scripts/requant_te_autoawq.py — same validated recipe as the ARK quant (bits 4, group 128, asym, iters 60, nsamples 32, seqlen 200, llava-mapped local calib data/ar_calib.json), but quantize_and_save(..., format="auto_awq").

Trim (the actual VRAM win)

scripts/trim_checkpoint.py — strips lm_head.* (1.16 GiB) and model.visual.* (1.07 GiB) from the checkpoint and rewrites shards + index. Result: 4.52 GiB slim dir (Qwen-Image-2.1-w4g128-autoawq-slim). Both modules are dead weight for T2I: the shim never calls lm_head, and text-only prompts never touch the vision tower. This beats quantizing embed_tokens (which autoround does not support anyway — no embedding-quant flag, and it would perturb every token’s input representation for ~0.8 GiB of system RAM only).

Component inventory (measured from safetensors headers)

Component bf16 quantized
TE layers 0-35 (w4 g128) 6.72 GiB 3.36 GiB
TE embed_tokens 1.16 GiB — (kept bf16)
TE lm_head 1.16 GiB trimmed
TE vision tower 1.07 GiB trimmed
DiT (fp8 weight-only) 13.25 GiB 6.63 GiB
VAE 1.26 GiB kept bf16

DiT → AWQ: investigation result

Known model-level caveat (not quant damage)

Exact prompt-text fidelity is loose at this config for every TE variant tried — same seed 42, prompt asking for a neon sign reading “QWEN IMAGE 2.1”: bf16 TE → no sign at all (city + orb); int8 TE → “OPEN”; w4 ARK → “HOTEL”; w4 AWQ-slim → “禁止停车 NO.001”. Coherent scenes, legible signage, but plausible-sign prior beats exact-string rendering. Semantic drift from the prompt (cabin → portrait) behaves the same across variants. No gray-noise/black outputs in any validated run; pixel stats (std 51–108) and vision checks recorded in the logs.

Reproduction

## env: conda qwen-image-21 (torch 2.14 cu130, diffusers 0.41.0.dev0, transformers 5.17,
##      autoround 0.15.1, autoawq 0.2.9)
## 1) requant TE to native AWQ (GPU, ~6 min quant + validation)
CUDA_VISIBLE_DEVICES=<free-gpu> python scripts/requant_te_autoawq.py
## 2) trim lm_head + vision tower
python scripts/trim_checkpoint.py \
  $WORKSTATION/models/qwen3vl-te-ar-w4a16/Qwen-Image-2.1-w4g128-autoawq-full \
  $WORKSTATION/models/qwen3vl-te-ar-w4a16/Qwen-Image-2.1-w4g128-autoawq-slim
## 3) all-resident generation benchmark (TE awq on GPU + fp8 DiT + VAE)
CUDA_VISIBLE_DEVICES=<free-gpu> python scripts/gpu_resident_test.py <slim-dir>
## quality gate
CUDA_VISIBLE_DEVICES=<free-gpu> python scripts/validate_autoawq.py

Weights are NOT included (Qwen Research License Agreement — non-commercial; the model weights and quant derivatives live in $WORKSTATION/models/qwen3vl-te-ar-w4a16/ on the workstation only). Scripts read the Telegram bot token from $WORKSTATION/.hermes/.env at runtime; no secrets are committed.

Evidence map

data/quality-suite-20260924.json defines eight fixed prompt/seed pairs covering portraits, hands, typography, crowded scenes, architecture, food, illustration, and macro textures. Each is generated at 1024×1024, 40 steps, true CFG 4 with the same AWQ text encoder and pinned model revision. FP8 remains the default; the candidate changes DiT weights to dynamic INT8 and enables compiled VAE decoding. One seed per category is a visual spot check, not proof of quality equivalence.

Run each command sequentially on an available GPU (0 was used for this suite):

PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 python scripts/gpu_resident_test.py \
  --prompt-suite data/quality-suite-20260924.json \
  --output-dir samples/quality-suite-20260924/fp8
PYTHONNOUSERSITE=1 CUDA_VISIBLE_DEVICES=0 python scripts/gpu_resident_test.py \
  --prompt-suite data/quality-suite-20260924.json \
  --dit-weights int8-dynamic --compile-vae \
  --output-dir samples/quality-suite-20260924/int8
python scripts/serve_comparison.py

Use the qwen-image-21 environment. The gallery binds to this workstation’s private address at private workstation service, includes the two original pairs, and exposes only listed images. It polls for completed results every 10 seconds; each run publishes images and warm timings atomically after generation. Native pixel zoom, linked scrolling, prompts, seeds, and original PNG links aid inspection. Reported generation times exclude per-prompt warmup, model loading and PNG saving; warmup times and compilation checks are retained in each run’s results.json. For a persistent gallery on this host, the checked-in web/qwen21-comparison-8769.service is linked and enabled with systemctl --user.

Completed results: eight images per mode; mean warm generation 93.352 s FP8 versus 46.219 s INT8 (2.020×), maximum allocated memory 17.7 versus 17.4 GiB. All 16 runs used 40 steps / 80 transformer calls and recorded zero measured compilations. Both batches used only physical GPU 0. Configs/results are beside the PNGs; pixel differences are in data/quality-suite-20260924-image-comparison.json (not perceptual quality scores). No outputs were rerolled or retouched.

Visual spot checks: both menus retain requested wording and prices; portraits, hands, food and fur remain coherent. Architecture details differ, and the fox’s pose/tail changes noticeably. Neither market scene clearly meets the exact requested headcount. These examples do not establish lossless quantization. The gallery passed Chrome checks at 1440×1100 and 390×844, including scene selection, next/previous, native zoom, linked/unlinked scrolling and no runtime errors. Seven CPU tests pass. FP8 remains the default.