Back to the experimentSupporting notebook

Identity LoRA v2

Source: comfyui/lucianax-loras-2026-09-22/docs/IDENTITY_FINETUNE_V2.md · revision 6550ead3945b

User-authorized sequence: generate eight high-quality image variations, train a new Qwen Image 2.1 LoRA first, then train an LTX-2.5 identity LoRA. Keep the same woman and pink swimwear, include both smiling and neutral expressions, and caption each image accurately.

Eight lossless frames were extracted at spaced timestamps from the highest-resolution original video (1280×720). An eight-reference Qwen edit produced duplicate people and was rejected. The accepted route uses one original reference per independently seeded edit. All eight accepted images are 1376×768, Euler 25 steps, CFG 1, with no old Qwen LoRA applied. They passed manual visual review for a single subject, facial appearance, garment consistency and visible pose/expression. This is visual curation, not a numerical identity guarantee.

The dataset contains exactly eight synthetic PNGs with individual natural-language captions. The identity token is lucianax; the old action phrase is not the new trigger. Each caption separately describes the visible pose, expression, pink halter bikini, medium framing and beach setting. No caption labels every expression as a smile, and no synthetic image is labelled as blowing a kiss. References, prompts, captions and media remain local and are not stored here.

Qwen was completed first. The selected adapter refines the existing v1 identity adapter for 240 additional steps, rank/alpha 32, learning rate 3e-5, at a 768 resolution bucket. Final smiling and neutral samples were reviewed, followed by an independent native ComfyUI text-to-image evaluation with no input-image conditioning. The installed lucianax_v2 profile points to this selected refinement, not the rejected fresh-base pilot.

LTX training completed 320 steps in 13 minutes 55 seconds, initialized from its own v1 Dev adapter: rank/alpha 32, 320 steps, learning rate 5e-5, 512 bucket, single-frame examples, with CPU layer offloading. It learns appearance from the same eight corrected-caption images. It does not learn new temporal motion from still images. Qwen weights are not transferred to LTX. The LTX training base is the non-distilled Dev ConvRot INT8 transformer; subsequent inference evaluation used Dev BF16 across two transformer GPUs.

Eight closely related, front-facing synthetic images remain a narrow dataset. They do not establish generalization to unseen facial views, lighting, clothing or long video chains. Evaluate the resulting adapters with both neutral and smiling prompts and new motion before replacing existing defaults. Keep v1 checkpoints for comparisons.

Local dataset: WORKSTATION/rank32_refine. Conda environment reused: WORKSTATION/ai-toolkit. Per-image provenance, hashes and logs are in the parent directory. The dataset contact sheet and PNG/caption archive were delivered through Hermes Telegram as messages 12734 and 12735.

Sources: Qwen Image 2.1, official ComfyUI edit template, LTX trainer documentation.

Status: Both refinements completed; two sequential BF16 video segments validated and visually reviewed.

Pilot and refinement

The conservative rank-16, 240-step Qwen pilot did not learn the face adequately. An independent ComfyUI test at step 160 confirmed the problem beyond the trainer previews. All 192 LoRA up matrices were nonzero, so this was not a zero-weight adapter. Some previews interpreted the garment-shape word as printed triangles. Revised captions now specify plain solid-color pink fabric, preserving the same eight image files and identity trigger.

Three bounded candidates use the three physical GPUs: resume the rank-16 pilot to 800 steps at 1e-4; train rank 32 fresh for up to 640 steps at 1e-4; and refine the already trained rank-32 v1 Qwen adapter for 240 steps at 3e-5 on the eight new images. Each uses a separate dataset/cache directory to avoid concurrent cache writes. Compare fixed-seed neutral and smiling outputs before choosing a checkpoint. Refinement reuses prior identity learning and is not a from-scratch run. The two weaker candidates were deliberately interrupted after checkpoint comparison. The refined candidate completed all 240 steps before LTX started.

Reuse and validation

The shared local profile registry now includes lucianax_v2, with the unchanged identity trigger lucianax. Its ComfyUI UI and API JSONs are generated from the same registry, and the CLI accepts this profile. The MCP profile-list function exposes it without requiring any disabled server registration to be enabled. Existing v1 profiles are preserved. Disabling the Qwen LoRA bypasses both the adapter node and automatic trigger insertion. The integration tests cover that behavior; 13 project tests passed.

The selected Qwen checkpoint has 384 tensors and 192 nonzero up matrices, all finite. Independent native ComfyUI generations tested a neutral expression and a hand-to-hair pose with no reference image conditioning. Visual review found substantially closer identity than the discarded fresh-base pilots and consistent plain pink clothing. This is a qualitative assessment, not a calibrated identity score.

Qwen final neutral output, smiling preview, and corrected caption dataset ZIP were delivered to the user as Telegram messages 12736, 12737, and 12738. The corrected ZIP supersedes the caption wording in message 12735. No checkpoint binaries are committed to this repository.

LTX BF16 evaluation

The non-distilled Dev BF16 transformer was split across GPUs 0/1; the text encoder and video/audio components used GPU2. All three power limits remained 245 W. The loader matched all 1,632 adapter targets (680 on GPU0, 952 on GPU1), materialized the patches, and measured nonzero weight differences on both GPUs. The adapter checksum is d1144faa83b2a49eda230bc30d40038958f4992f307cd7db265acfb51d3802ca.

Two sequential 3-second clips were rendered at 1024×576, 24 fps, 30 Euler steps. The first starts from an original reference frame; the second starts from the preceding clip’s last frame. The actions are a gentle shoulder turn and a hand-to-hair gesture. Both passed full decode, duration, codec, audio, nonblack-content and gallery-provenance checks. Boundary SSIM was 0.971904; that measures the join, not facial identity.

Clip ComfyUI execution Seconds of execution per video second
Shoulder turn 511.273 s 170.424 s
Hair gesture 507.200 s 169.067 s

These totals include model/component work and saving; they exclude subsequent gallery archival waits. Combined execution was 1,018.473 s for 6 s of video, averaging 169.746 s per video second. This is a different resolution from the earlier 576×768 diagnostic, so it is not a controlled speed comparison against v1 or Distilled. Peak GPU memory was 21,576 / 22,650 / 23,456 MiB on GPUs 0/1/2 respectively.

Visual inspection at half-second intervals found broadly consistent facial appearance, blue eyes, dark hair and plain pink clothing across the pair, without an obvious radical person replacement or forced kiss gesture. This remains a qualitative six-second check; it does not establish identity robustness for 30 seconds or unseen viewpoints.

The exact API graphs are in WORKSTATION/dev-identity-v2. They use the local sharding and verified-adapter extensions around the Dev graph, and are not an unmodified upstream workflow or editable UI export. The Qwen v2 UI workflow is installed separately in the ComfyUI workflow library. Shared skill links for Codex, OpenCode, Pi, OMP and Hermes all resolve to the updated project skill.

Final delivery

The joined six-second MP4 passed the same full decode, audio, duration, content and strict gallery-provenance gates. It was delivered as Telegram message 12742; the comparison sheet as 12743; and the final Spanish report as 12744. Training previews were messages 12739–12740 and the Qwen comparison was 12741. Project implementation commit: 3a1faca. After completion, ComfyUI unloaded the models; all three power limits remained 245 W.