This guide covers H3 reference-to-video training: choosing a training goal, preparing reference material and target images or videos, and creating a job in AI Toolkit. The video demonstrates both complex video concepts and anime-to-live-action editing data. These require different dataset preparation.

It accompanies the September 18 video tutorial, using the cloud AI Toolkit environment provided with that video.

1. Decide what the LoRA should learn

GoalMain materialsWhat to check
A characterCharacter imagesConsistent visual identity
An art styleVaried images in the same styleMore than one character, pose and setting
An editing transformationSource material + target materialEach input must match its target
A complex video conceptReferences + target video + captionCorrect reference pairing, frame counts and action descriptions

H3 supports different conditioning modes. Text-to-video and first/last-frame tasks use different models and training modes from tasks with arbitrary references. This tutorial uses Ref2V. Do not reuse a previous text-to-video or image-to-video configuration by changing only the dataset path.

If you only want to learn an anime show’s style from images, see Train an H3 Anime-Style LoRA with Musubi Tuner.

2. Video data: organize frame counts and references first

The complex-concept dataset in the video is grouped by frame count, with corresponding reference images for each target video. After uploading, each bucket has a clear video length rather than leaving you to guess during training.

H3 video uses the 17k+5 frame-count rule. Examples in the video include 22, 73, 141 and 158 frames; the author’s H3 branch documentation also confirms this grid. Unsupported frame counts can cause training errors.

The demonstrated AI Toolkit build provides auto frame counts, which rounds down to a supported frame count. It handles frame alignment, but it cannot verify reference-image pairings or whether captions accurately describe the video.

Check these three things when captioning:

  • Does the caption describe this particular clip? Batch AI captioning can mix clips up.
  • Are the actions, people and scene details actually present? Remove invented details.
  • Do the reference images correspond to the correct target video?

The video recommends using official prompt guidelines to organize descriptions of complex concepts, while still spot-checking the results. A simple image concept does not require the entire complex-video captioning process.

3. Editing data: control is the input; targets are the desired output

Another example converts anime material into a live-action look. It uses two matched sets:

  • control: original inputs, such as anime images.
  • targets: desired outputs, such as the corresponding live-action images.

The author separates target styles into groups: one is more natural, while another has a stronger beauty-filter look. Training increasingly reflects those dataset choices. Inspect the targets before deciding why the model learned a particular look.

This image-editing example does not use a separate caption for each image. Instead, a shared default caption describes the transformation. If you already have per-image captions, use them. The absence of sidecar text files does not mean there is no text conditioning.

4. Create the training job

After uploading the data, create a job in AI Toolkit or copy the tutorial’s existing job and adjust it. Check the following in order:

  1. Model and mode: select the H3 reference-to-video model you intend to train.
  2. Training method and adapter: confirm whether the job uses a training adapter or another training method.
  3. Dataset binding: bind original material to control and target material to targets.
  4. Text conditioning: enter a shared transformation description or use your prepared individual captions.
  5. Sampling: replace the example paths and prompts with your own test material.

Image and video jobs also use different settings. The image-only example has no audio and uses a frame count of 1. Do not copy those image settings unchanged into a video job. Select the appropriate input type for reference images versus reference videos.

5. Do not combine incompatible training methods

The video compares adapters and training methods, including methods that use a teacher model to preserve the base model’s capabilities. These options change training behavior and runtime; enabling more is not automatically better.

The circlestone-labs image training adapter has an explicit restriction: do not enable CFG-enhanced training or guidance preservation loss alongside it. Follow the author’s restriction instead of turning on every option that sounds quality-related.

For the first run, use one internally compatible configuration from the tutorial and verify the data and sampling. When changing adapters, recheck the training method rather than replacing only a file path.

6. Rank, steps and VRAM settings

SettingExample in the videoHow to interpret it
LoRA rank16 or 32Candidate values to compare for your task
Training steps5000A training plan, not a guarantee that the final checkpoint is best
Checkpoints to retain10Keep intermediate results for comparison
OffloadingAdjust for your GPUTrade transfer time for lower VRAM use
ResolutionReduce when VRAM is insufficientImages, videos and reference inputs all affect memory use

Do not use the demonstrated H20 iteration speed as an estimate for your machine. Run caching, sampling and a short training job with your own data first, then extend the job once resource use is acceptable.

7. Download the LoRA and check it against your goal

Watch the samples during training and download the resulting LoRA. For editing, check whether the transformation works while preserving the intended character, pose and content. For complex video concepts, also check the learned actions and reference constraints.

If the result increasingly develops an unwanted beauty-filter look or another fixed style, revisit the target material. The video shows training results following those dataset tendencies.

For another paired-editing example, see Qwen Image 2.1 Editing LoRA: Anime to Live Action. The dataset organization is comparable, but the model and training settings are not interchangeable.

Resources