The follow-up checks use three action categories for each trained concept:
- turn around;
- dance;
- pull hair back.
The exact prompts are excluded by repository policy. LTX prompts retain the training trigger and append one requested motion. Qwen prompts retain only the appropriate identity token and append the same action category.
LTX-2.5 results
| Action | Seed | Output properties | Visual observation | Telegram |
|---|---|---|---|---|
| turn around | 3162030 | 512×320, 73 frames, 24 FPS, H.264 + AAC | front view progresses through profile to a back-facing pose | 12718 |
| dance | 3162031 | 512×320, 73 frames, 24 FPS, H.264 + AAC | clear dance motion with a raised-arm finish | 12719 |
| pull hair back | 3162032 | 512×320, 73 frames, 24 FPS, H.264 + AAC | both hands move behind the head and hold the hair | 12720 |
All three files passed a complete FFmpeg decode. Inference used two full transformer replicas, one on GPU 0 and one on GPU 1, with CFG work splitting. The text encoder and video/audio VAEs ran on GPU 2. Peak sampled memory was 21,764 MiB, 19,868 MiB, and 2,764 MiB respectively; all three cards retained the 245 W cap.
Qwen-Image 2.1 results
| Concept | Action | Seed | Output properties | Visual observation | Telegram |
|---|---|---|---|---|---|
lucianax |
turn around | 3162033 | 768×1024 JPEG | back-facing beach pose with an over-the-shoulder look | 12721 |
lucianax |
dance | 3162034 | 768×1024 JPEG | clear, energetic dance pose | 12722 |
lucianax |
pull hair back | 3162035 | 768×1024 JPEG | both hands behind the head holding the hair | 12723 |
lucianat |
turn around | 3162036 | 768×1024 JPEG | back-facing indoor pose in the learned dark styling | 12724 |
lucianat |
dance | 3162037 | 768×1024 JPEG | dynamic dance pose in the learned dark styling | 12725 |
lucianat |
pull hair back | 3162038 | 768×1024 JPEG | hands behind the head pulling back the ponytail | 12726 |
The Qwen inference harness reopened each completed job at step 800 and performed sampling without optimizer steps. It re-serialized both final checkpoints; their final hashes and unchanged training metadata are recorded in the structured evidence.
Interpretation
Every requested action was legible in its output. The strongest limitation remains dataset entanglement: the LTX result preserves the beach scene and wardrobe because all three source clips depicted the same scene. Multiple seeds, varied camera positions, and held-out identity comparisons would be needed to measure generalization more rigorously.