The user observed a substantial identity change after the first clip of the sequential video.
The official Distilled workflow and the first Dev BF16 benchmark used a Qwen-generated initial
image but did not apply the trained LTX identity LoRA. The continuity API’s use_lora flag
controls Qwen image generation only. A last-frame continuation and high boundary SSIM do not
establish preservation of facial identity.
The no-LoRA Dev BF16 chain was stopped during its second segment, and its automatic delivery
watcher was terminated. The completed first segment remains a no-LoRA control. The interrupted
run is marked STOPPED_IDENTITY_CORRECTION; it is not a completed 30-second benchmark.
Verified correction
The separate Dev route now inserts LucianaVerifiedDevIdentity between the BF16 transformer
loader and all guidance/sampling branches. It loads the Dev-trained step-600 adapter at strength
1.0, verifies its SHA-256, requires every expected adapter target to match, and materializes the
patches before sampling. Both the initial and continuation graphs include the training trigger.
The official Distilled route is unchanged.
Adapter SHA-256: 5ce4ece9e8ec8ccd8cd3aebd6104c62c18b42205bd8036ec9b416881d13cc810.
The live preflight matched all 1,632 adapted weights: 680 on GPU 0 and 952 on GPU 1.
Sampled materialized weights differed from their saved base weights on both GPUs (maximum
absolute differences 0.0009765625 and 0.00074005126953125). This establishes actual LoRA
application, not a guarantee of visual identity.
The first loading attempt exhausted GPU memory while replacing parameters retained in the sharding loader’s load list. The correction uses the existing ModelPatcher in-place update option and CPU backups, preserving reversible base weights without duplicate GPU parameters. The repeated live preflight passed. All three GPUs retain 245 W power caps.
Visual validation
Two sequential diagnostic clips completed at 576×768, 72 output frames each, 24 FPS, 30 Euler steps. These are three-second identity probes, not a replacement for the requested full-resolution timing comparison. The transformer is BF16 across GPUs 0/1; encoders and VAEs use GPU 2. Each continuation consumes the preceding final frame. The original source image remains the fixed visual evaluation reference; it is not separately injected into every continuation. No untested claim of independent identity reference conditioning is made.
Dataset limitations remain: three seeds of one scene do not cover diverse views, lighting, wardrobe, or motion. The adapter may entangle identity with the training action and background. Inspect faces at the beginning, middle, and end of each segment against the fixed reference; reject drifting segments before extending the chain. Increase adapter strength only through controlled tests; it can also exaggerate training-scene bias. Improve the training dataset if identity remains unstable despite verified application.
Local evidence is under WORKSTATION/dev-bf16-identity and
WORKSTATION/lora-audit.json. Media, reference images,
model files, and exact prompts are excluded from this repository.
The official LTX LoRA guide describes adapters as improving subject fidelity and recommends prompts aligned with the training data. It does not establish a guarantee that an adapter preserves identity for every motion or arbitrarily long sequence.
Results
Both three-second clips passed full video/audio decode, content validation, and strict gallery provenance. The boundary SSIM was 0.979736; this measures the overlap frame, not identity. The initial short clip took 368.340 seconds and its continuation 395.136 seconds of server execution: 12 min 43 s total for six generated seconds, about 127.25 compute seconds per output second at this diagnostic resolution. Peak GPU memory was 21,116 / 22,266 / 23,456 MiB.
Visual inspection of middle/end frames found broadly consistent hair, eyes, face appearance, wardrobe and scene across these two short clips, without the previous abrupt redesign. Facial appearance is not exact; pose and expression change. The kiss gesture reappears while the second clip moves its hand to the hair, consistent with action/trigger entanglement. This is not evidence of robust identity over 30 seconds or unseen motions. The original source still and a training frame accompany the diagnostic contact sheet for direct human comparison.
The overlap checker initially used the canonical five-second frame index. The probe now checks frame 71 against frame 0 for these three-second clips; saved renders were reused, not regenerated. The ComfyUI worker was unavailable after the accepted clips completed, with no captured error explaining its exit. It was restarted for final concatenation publication. No completed render was lost.
The earlier full-resolution no-LoRA controls remain separate: Distilled 62.812 seconds versus Dev BF16 1,116.402 seconds for the first five-second clip (about 17.8× slower for Dev). Both lacked the LTX identity adapter. Do not compare these directly with the smaller identity probes to infer the speed cost of LoRA. The 30-second Dev comparison was interrupted, not completed.
Delivered via Hermes Telegram: video 12731, reference contact sheet 12732, report 12733.
Implementation and exact API graphs are committed in luciana-media-mcp at c74f263.