To give H3-generated videos the visual style of an anime show, you can start by training a style LoRA on images. Using the video’s Frieren example, this guide covers image selection, captions, the training adapter, low-VRAM settings, caching and sample checks.
It accompanies the September 25 video tutorial. The demonstration uses an H3-enabled Musubi Tuner branch and its accompanying cloud image. The displayed model paths and presets belong to that environment; use your own paths for a local installation.
What to prepare
| Item | What you need |
|---|---|
| Training tool | The Musubi Tuner H3 branch used in the video |
| Models | The H3 model, text encoder and VAE appropriate for the training task |
| Dataset | Clean anime images with corresponding captions |
| Training adapter | The circlestone-labs H3 image training adapter |
| Test content | Sampling prompts that let you compare style, characters and motion |
The video demonstrates configuration changes on a 4090 and also uses an H20 for full training. The author prefers having more VRAM available. On a 24GB GPU, the quantization, resolution and swap settings below are especially relevant.
1. Represent the style, not just one character
The dataset in the video contains about 400 images; the author’s suggested range is 350–500. More important than the count is whether the images represent the show’s visual style.
Include a variety of:
- Characters and character combinations.
- Close-ups, medium shots and wide shots.
- Poses, expressions and camera angles.
- Interiors, exteriors and shots without characters.
After extracting frames, remove near-duplicate consecutive frames and distractions such as subtitles. If almost every image has the same character and composition, the training can entangle that character with the style and generalize poorly to other scenes.
Downloaded screenshot collections also need curation. The author spends substantial time on this step. Having a folder of images is not the same as having a finished dataset.
2. Captions: trigger word first, image content next
Each image in this tutorial has a caption beginning with the style LoRA’s trigger word, followed by the visible characters, actions and setting.
When several characters appear, distinguish them clearly. The examples describe characters individually so the model can associate their names with their appearances.
For this anime-style training, the author does not repeatedly add a generic “anime style” phrase to every caption. Instead, captions describe specific content while the shared visual style comes from the dataset. This is the captioning approach used in this example.
Spot-check image/caption pairs for swapped character names, incorrect person counts and invented details.
3. What the image training adapter does
H3 ultimately generates video, but this job learns a style from images. Training only on images can also change the model’s video behavior, weakening motion or producing a slideshow-like result.
The circlestone-labs adapter is designed for this training approach. According to the author’s instructions, merge it into the base model before training the LoRA. When using the trained LoRA for inference, omit the training adapter.
Do not combine it with CFG-enhanced training or guidance preservation loss. These combinations have restrictions; do not enable them simply because the trainer offers switches. The author also notes that video and mixed image/video training have received less testing.
4. Adjust the preset to fit your VRAM
The demonstrated default preset was made for an H20 and produces a VRAM warning on a 4090. Check the dataset, caching and training stages separately.
Dataset and resolution
Place the images and captions in your dataset directory and point the job to it. Replace the example paths.
If VRAM is tight, lower the training resolution first. The video demonstrates moving toward 768 or 1024. Get caching and sampling working before increasing the image size.
Text encoder and training precision
Text-encoder caching and model training are separate stages. Fitting one stage into VRAM does not guarantee that the other will fit.
In the 24GB example, the author switches to a low-VRAM text-encoder configuration and uses INT8-related options for training. The format name in the subtitles was checked against the author’s branch documentation and identified as NVFP4; it is not a model name to reconstruct from a transcription error. See the H3 branch documentation for the actual formats and corresponding files.
Block swap
Block swap moves some model weights between system RAM and VRAM. It reduces VRAM use for weights but adds transfer overhead. After checking the memory estimate, the author uses a swap count of 20 in the video. This is not a universal value for every dataset.
Check the estimated memory use for your models and inputs before changing the swap count. If high resolution or long sequences are the main source of pressure, reduce those as well instead of relying only on more swapping.
5. Cache in order, then train
Follow the sequence shown in the video:
- Cache latents: process the latent representations of the training images.
- Cache text-encoder outputs: process the captions used for conditioning.
- Start training: use the caches from the first two stages.
If a stage fails, resolve that stage’s path or VRAM problem before continuing. The demonstrated interface supports both separate stages and a combined start action.
6. Check sampling at the start
Generating samples uses additional VRAM. If the first sampling attempt happens only at an intermediate checkpoint, training may appear to work and then fail during sampling.
When you need training samples, enable initial sampling and verify that it completes. Use your own test prompts, dimensions, frame counts and seeds rather than leaving someone else’s preset test content in place.
If you do not need samples during training, you can disable sampling and compare checkpoints in ComfyUI afterward.
7. The final checkpoint is not always the best
The video configures 20000 steps, but the author finds the result usable at around 16000 steps and observes weaker motion after further training. Compare both style and movement, rather than judging only whether a still frame increasingly resembles the source show.
Keep intermediate versions and compare them with the same prompts. A short startup demonstration using very few images is not a reliable estimate of full-dataset training time.
For reference-image, video-editing or complex-concept training rather than anime style alone, continue with Train H3 Reference-to-Video and Editing LoRAs with AI Toolkit.