Journal

The faster image pipeline that produced nothing

The balanced two-GPU path took 48.1 seconds but every image was invalid flat gray noise.

Vlad / experimentos.
The boundary

The sharding mechanism is a suspect, not a proven cause. The broken outputs were removed; the failure record remains.

The question

The two-GPU path appeared faster until the outputs were inspected. Every image was invalid, so the experiment records a failure and an unresolved sharding suspect rather than promoting the shorter timing.

From the original notebook

Status: BROKEN / DO NOT USE

2026-09-21: the broken PNGs and .broken marker files listed below were deleted at the user’s request; only this record remains.

Source project used to generate these

  • Model: Qwen/Qwen-Image-2.1 (Qwen-Image-2.1)
  • Pipeline: diffusers.QwenImage21Pipeline
  • Script: WORKSPACE/temporary-artifact
  • Config: torch_dtype=bfloat16, device_map="balanced" across 2× RTX 3090
  • Run: CUDA_VISIBLE_DEVICES=1,2 conda run -n qwen-image-21 python WORKSPACE/temporary-artifact
  • Steps/size: 40 steps, 1024×1024, seeds from random.Random(2026)

Failure

Every image is flat gray noise (no content). The device_map="balanced" sharding of the diffusers pipeline is the suspect: latents/tensors split across devices, so the denoise loop never produces a coherent image. Speed was fine (48.1 s, 1.42× vs 68.3 s single-GPU baseline) but output is invalid — the 2-GPU path is not usable as-is.

Files

file md5 (short) verdict
qwen21_2gpu_1.png 53cbc3167ff5 BROKEN (flat gray noise)
qwen21_2gpu_2.png e4072fe100ea BROKEN (flat gray noise)
qwen21_2gpu_3.png b5ce06d363dd BROKEN (flat gray noise)
qwen21_2gpu_4.png 7fe1bed571c0 BROKEN (flat gray noise)
qwen21_2gpu_5.png 1214027f06ea BROKEN (flat gray noise)

Working alternatives (same model): single-GPU bf16 + enable_model_cpu_offload() (WORKSPACE/temporary-artifact) and fp8 weight-only DiT + offload (WORKSPACE/temporary-artifact).

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS