Back to the experimentSupporting notebook

Photo15x: Qwen Image 2.1 photographic-style LoRA

Source: comfyui/photo15x-qwen21-2026-09-22/README.md · revision 6550ead3945b

The user requested fifteen randomly selected, high-quality photographic-looking images from ComfyUI input, without duplicates, and a new Qwen Image 2.1 LoRA. Inspection found 213 direct image files, reduced to 97 unique decoded images after excluding 116 exact copies. Source files were not deleted or altered.

Visual curation excluded illustrations, fantasy scenes, explicit images, obvious face-swap files, repeated scenes and poor/small images. Twenty-one photographic portrait candidates remained. The twenty largest by source file size formed the random pool; Python random.Random(2026092215).sample(pool, 15) selected the final dataset. Dimensions and visible detail were also checked: PNG versus JPEG compression makes raw file size an imperfect quality metric. All fifteen selected images are distinct, with minimum pHash distance 22 and manual crop/scene review. Some source names indicate generated or edited provenance; this is not a claim of fifteen verified camera originals.

The dataset contains multiple people, so the run is explicitly a photographic-style adapter using trigger photo15x, not a new single-person identity or a Lucianax update. An optional clarification about identity versus style was presented; absent a contrary instruction, the mixed-subject style interpretation was stated before training began.

Each image has a descriptive natural-language caption of the visible subjects, clothes, composition and lighting, including the literal trigger because text embeddings are cached. Originals were EXIF-normalized into lossless PNGs without AI enhancement. Exact captions and images remain local under WORKSTATION/dataset.

Training starts from the Qwen Image 2.1 base, without any existing character adapter: rank/alpha 32, 600 steps, LR1e-4, AdamW8bit, BF16 LoRA, official INT8 ConvRot base, 768 resolution buckets, cached latents/text, transformer on GPU0, text encoder offloaded then unloaded. The existing ai-toolkit Conda environment is reused. Baseline samples and checkpoints every 200 steps support comparison on fixed prompts/seeds.

Status: DONE. The trainer exited successfully after 600 steps (23m27s training loop, plus model loading/caching and final samples). Checkpoint 400 was selected: the 600-step fixed-seed outputs introduced unrequested garment cutouts and heavy makeup. This is a small visual comparison, not evidence of universal improvement over base Qwen.

Native ComfyUI evaluation compared base, 200, 400, and 600 with identical prompts/seeds, 1024x1024, 25 Euler steps, CFG1. Each warm image took about 16 seconds on GPU2. An additional checkpoint400 test at strength0.5 gave a milder effect. Default workflow strength remains1.0; reduce to0.5 when the style dominates clothing/background details. All384 selected tensors are finite and all192 LoRA B matrices are nonzero.

Installed model: WORKSTATION/photo15x_qwen_image_21_lora_v1.safetensors points to the retained step400 checkpoint. Step200 and final600 remain in ai-toolkit output. Profile photo15x is in the shared CLI/MCP registry; no disabled MCP was enabled. Editable ComfyUI JSON: WORKSTATION/Qwen Image 2.1 - photo15x.json. Shared agent skill updated; existing project suite:13 passed.

Local comparison: WORKSTATION/comparison.jpg; selection and tensor audits are alongside it. Source images, exact captions/prompts and model binaries stay outside this repository. Three GPU power limits remained245W. Models were unloaded after review.