Journal

Character LoRAs: identity, motion, and continuity

Three initial adapters completed: 600 LTX steps and two Qwen identity runs of 800 steps each.

Vlad / experimentos.
The boundary

Training completion and playable files do not establish identity or action fidelity. Source and generated media remain private.

The question

Character adapters need more than completed training. This project records the initial identity and motion runs, later action probes, continuity work and identity corrections, while leaving private reference media outside the article.

From the original notebook

This experiment records three character LoRA training runs and their local inference checks on three RTX 3090 GPUs. Model weights, source media, generated media, captions, and verbatim prompts remain outside this repository.

Question

Can three same-scene videos teach LTX-2.5 a character and a kiss action, and can still-image datasets teach Qwen-Image 2.1 the related lucianax and lucianat identities?

Training runs

Model Concept Input Steps Rank / alpha Final artifact
LTX-2.5 dev character plus kiss action 3 videos, 121 frames per training clip 600 32 / 32 674,249,600-byte Safetensors, 3,264 tensors
Qwen-Image 2.1 lucianax identity 20 frames extracted from the highest-resolution source video 800 32 / 32 159,436,416-byte Safetensors, 384 tensors
Qwen-Image 2.1 lucianat identity 22 ChatGPT images 800 32 / 32 159,436,416-byte Safetensors, 384 tensors

All three jobs used BF16 LoRA saves, ConvRot INT8 quantized base models and text encoders, cached latents and text embeddings, AdamW 8-bit, batch size 1, and gradient checkpointing. Each final Safetensors file was opened with safetensors.safe_open; tensor counts and training metadata were read successfully.

LTX inference topology

The post-training video check used ComfyUI MultiGPU_WorkUnits after applying the learned LoRA:

flowchart LR
    M[LTX-2.5 dev INT8] --> L[Character LoRA]
    L --> G0[Transformer replica: GPU 0]
    L --> G1[Transformer replica: GPU 1]
    C[CFG 3 video / CFG 7 audio] --> G0
    C --> G1
    T[Gemma encoder: GPU 2] --> C
    V[Video and audio VAEs: GPU 2] --> O[H.264 + AAC output]
    G0 --> O
    G1 --> O

This is full transformer replication with CFG work splitting. It is not tensor or layer sharding. The three initial probes used 25 Euler steps, 512×320, 49 frames, and 24 FPS. All three completed in ComfyUI, decoded fully with FFmpeg, and contained H.264 video plus AAC audio. Peak sampled GPU memory was 21,690 MiB on GPU 0, 19,794 MiB on GPU 1, and 2,540 MiB on GPU 2. All cards were capped at 245 W.

Initial observations

  • The LTX samples consistently reproduced a brunette woman, the learned beach context, and the kiss gesture.
  • The first Qwen lucianax sample produced a coherent brunette beach portrait with the learned gesture.
  • The first Qwen lucianat sample produced a coherent brunette editorial portrait.
  • The LTX concept is visibly entangled with the source scene and wardrobe. The action probes in this experiment test whether unseen motions override that learned context.

Action-probe observations

The three LTX follow-up videos all followed the requested motion while retaining a consistent brunette subject and much of the learned beach context. The turn-around sequence progressed from front view through profile to a back view; the dance sequence showed clear body and arm movement; and the hair sequence ended with both hands behind the head holding the hair. Each output contained 73 frames at 512×320 and 24 FPS, plus AAC audio, and passed a full FFmpeg decode.

Both Qwen LoRAs also followed all three action categories in 768×1024 still images. lucianax retained the beach-oriented appearance, while lucianat retained the darker editorial styling. The turn-around images used a back or over-the-shoulder pose, the dance images showed a dynamic pose, and the hair images placed both hands at or behind the head.

The action tests therefore show usable motion and pose control, although the LTX identity remains coupled to the narrow scene and wardrobe represented by its three training videos. This is a qualitative result from one seed per action, not an identity benchmark.

Validation

  • AI Toolkit targeted tests: 2 passed.
  • LTX final checkpoint metadata: step 600, epoch 199.
  • Qwen lucianax checkpoint metadata: step 800, epoch 19.
  • Qwen lucianat checkpoint metadata: step 800, epoch 18.
  • Every reported MP4 completed a full ffmpeg -v error -i INPUT -f null - decode.
  • Three LTX action videos and six Qwen action images were delivered through Telegram; message IDs are included in the structured evidence.

Measured values and hashes are in data/training-summary.json. Action probe outcomes are recorded in docs/ACTION_PROBES.md.

Reproduction boundary

The source media, generated media, prompts, captions, and model weights are intentionally excluded. Reproduction requires the local AI Toolkit and ComfyUI checkouts, the listed base models, and the private source datasets. The report distinguishes observed outputs from qualitative interpretation.

Reusable ComfyUI and agent integration

See ComfyUI workflows, MCP, and official LTX migration for the installed Qwen JSON workflows, five-agent integration, and validated official LTX-2.5 I2V results.

The 30-second continuation test reuses the first official clip and adds five sequential last-frame-conditioned actions, with measured speed and boundary validation.

Reusable 30-second workflow

ComfyUI and MCP integration includes the four JSON variants, durable agent tools, optional Qwen LoRA, and fresh end-to-end validation. Download workflow ZIP.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (8)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS