This tutorial trains an editing LoRA on paired anime originals and live-action target images, teaching Qwen Image 2.1 a specific transformation. The process is to prepare paired images, configure the job, test sampling first, and select a LoRA by comparing intermediate samples.

It accompanies the September 22 video tutorial. The instructions follow the video’s cloud AI Toolkit environment, which already contains the models. For local installation, use the project documentation linked at the end.

What to prepare

  • An AI Toolkit environment with Qwen Image 2.1 support and the appropriate models.
  • Original images and a corresponding target image for each one.
  • A trigger word or phrase describing the desired transformation. The video uses examples such as photo realistic and Live action.

Decide which transformation you want to learn before collecting data. For a live-action target, choose images with the specific photographic qualities you want. Mixing visibly different target looks does not tell the model which one you prefer.

1. Pair each original with its target

The two image groups have different roles in AI Toolkit:

FolderContentsExample in this tutorial
controlInputs before editingOriginal anime images
targetsOutputs the model should learn to produceCorresponding live-action images

The filenames in each pair must match. The model needs to learn how this particular input should change, not an arbitrary relationship between unrelated images.

You can organize the folders as follows, replacing the example filenames with your own:

dataset/
├── control/
│   ├── character_a.png
│   └── character_b.png
└── targets/
    ├── character_a.png
    └── character_b.png

Upload both folders into the dataset directory and refresh AI Toolkit’s dataset page. Open sample pairs to check that they match, that inputs and targets are not reversed, and that the target look is the one you actually want to learn.

2. Create an editing training job

If you use an existing job provided with the tutorial, copy it before making changes. This preserves the original configuration and makes precision or parameter comparisons easier.

Configure the following:

  1. Select Qwen Image 2.1 and confirm it is the model you intend to train.
  2. Bind target images to target and original images to control in the dataset settings.
  3. Enter your trigger phrase and use the same transformation description in test sampling.
  4. Replace the sample input images with your own originals. Otherwise, the samples only show performance on someone else’s examples.

The demonstrated editing configuration uses shared default text conditioning and a trigger phrase rather than individual captions for every image. Not captioning each image separately does not mean prompts are unnecessary: the model still needs a description of the intended edit.

3. Choose a balance of precision, VRAM and speed

The video compares two approaches:

ApproachObservation in the videoWhat to consider
bf16Demonstrated on a 4090; the author prefers the training resultAvailable VRAM and sample quality
INT8 / 8bit quantizationA 3090 setup is discussed; faster and less demanding on VRAM, but worse in this testWhether the saved resources justify the observed quality loss

These observations apply to the demonstrated models, data and configuration. GPU model alone does not establish whether a job will fit. In particular, distinguish training memory use from sampling memory use: a job can start training successfully and still run out of VRAM when generating samples.

The video uses a resolution of 1280 and adjusts offloading when VRAM is insufficient. With your own data, begin at a resolution that samples reliably, then increase it gradually.

4. Test sampling before a long training run

Watch the initial sampling output when you start the job. It helps check data bindings, prompt compatibility and whether the model can generate with the current settings.

If sampling runs out of VRAM, troubleshoot in this order:

  1. Lower training/sampling resolution. Do not insist on high resolution when the current setup already exceeds memory.
  2. Check offloading. One demonstrated failure occurs with offloading disabled.
  3. Run sampling again. Continue to full training only after the problem is resolved.
  4. If it still cannot run, consider quantization and compare the quality again.

The video adjusts an offloading-related setting to 30, but the subtitles do not preserve the exact field name. Do not interpret this as instructions to set any percentage field to 30. See the operation at about 5 minutes 30 seconds.

5. How many steps, and which checkpoint?

The job in the video reaches 5000 steps, but the author sees changes at 500 steps and good results around 1000 and 1750. This dataset does not necessarily need the full planned run.

Keep intermediate checkpoints and compare them with the same input images and prompts:

  • Has the target style appeared?
  • Are the character features and composition that should remain intact still preserved?
  • Does further training actually improve on the previous version?

One iteration in the video takes about 5.44 seconds; “around two hours” describes the scale of that particular run. Your timing depends on the step count, resolution, quantization and offloading settings.

Use the trained LoRA

Download the LoRA and test it in a matching Qwen Image 2.1 editing workflow. You can also use AI Toolkit’s test interface: select the model and LoRA, generate a result, and confirm the effect before moving to your usual ComfyUI workflow.

For the base model’s generation, editing and transparency features, see Qwen Image 2.1: Generation, Editing and Transparent Images.

Resources