Back to the experimentSupporting notebook

Action probes

Source: comfyui/lucianax-loras-2026-09-22/docs/ACTION_PROBES.md · revision 6550ead3945b

The follow-up checks use three action categories for each trained concept:

  1. turn around;
  2. dance;
  3. pull hair back.

The exact prompts are excluded by repository policy. LTX prompts retain the training trigger and append one requested motion. Qwen prompts retain only the appropriate identity token and append the same action category.

LTX-2.5 results

Action Seed Output properties Visual observation Telegram
turn around 3162030 512×320, 73 frames, 24 FPS, H.264 + AAC front view progresses through profile to a back-facing pose 12718
dance 3162031 512×320, 73 frames, 24 FPS, H.264 + AAC clear dance motion with a raised-arm finish 12719
pull hair back 3162032 512×320, 73 frames, 24 FPS, H.264 + AAC both hands move behind the head and hold the hair 12720

All three files passed a complete FFmpeg decode. Inference used two full transformer replicas, one on GPU 0 and one on GPU 1, with CFG work splitting. The text encoder and video/audio VAEs ran on GPU 2. Peak sampled memory was 21,764 MiB, 19,868 MiB, and 2,764 MiB respectively; all three cards retained the 245 W cap.

Qwen-Image 2.1 results

Concept Action Seed Output properties Visual observation Telegram
lucianax turn around 3162033 768×1024 JPEG back-facing beach pose with an over-the-shoulder look 12721
lucianax dance 3162034 768×1024 JPEG clear, energetic dance pose 12722
lucianax pull hair back 3162035 768×1024 JPEG both hands behind the head holding the hair 12723
lucianat turn around 3162036 768×1024 JPEG back-facing indoor pose in the learned dark styling 12724
lucianat dance 3162037 768×1024 JPEG dynamic dance pose in the learned dark styling 12725
lucianat pull hair back 3162038 768×1024 JPEG hands behind the head pulling back the ponytail 12726

The Qwen inference harness reopened each completed job at step 800 and performed sampling without optimizer steps. It re-serialized both final checkpoints; their final hashes and unchanged training metadata are recorded in the structured evidence.

Interpretation

Every requested action was legible in its output. The strongest limitation remains dataset entanglement: the LTX result preserves the beach scene and wardrobe because all three source clips depicted the same scene. Multiple seeds, varied camera positions, and held-out identity comparisons would be needed to measure generalization more rigorously.