Qwen Image 2.1 can generate images from text and edit existing images. Based on the video’s two ComfyUI workflows, this guide explains prompt rewriting, sampling steps, output dimensions and the resolution budget, along with practical observations from the tests.

It accompanies the September 21 video tutorial. Start by getting text-to-image generation working, then try editing and transparent images so problems are easier to isolate.

Choose a workflow

What you want to doWorkflowMain inputs
Design a scene, poster or illustration from scratchText-to-imageA description of the idea
Modify an image, create character views or color artworkImage editingOriginal image + editing instructions
Generate transparent assets or change an existing backgroundChoose based on whether you already have an input imageTransparency instructions and an optional input image

The video uses the Qwen Image 2.1 generation model, a Qwen3-VL text encoder and the corresponding VAE. These components must match the workflow. Replacing only the generation model does not ensure the other components are compatible.

1. Text-to-image: describe the idea, then decide whether to rewrite the prompt

The text-to-image workflow includes the official prompt-rewriting model. It expands a short idea into a more detailed visual description and passes that description to the generation model.

Use it in this order:

  1. Choose the image dimensions and resolution.
  2. Enter the subject, scene or visual idea.
  3. Use prompt rewriting if the description needs more detail.
  4. Generate an image from the expanded description, review the result, and adjust the original idea.

If your prompt is already clear, check whether the rewritten version preserves your intent rather than judging it by length alone. Separate rewriting models are available for generation and editing; see Prompt Rewriting.

2. Sampling steps and precision

These are the choices demonstrated in the video. Compare results in your own workflow:

SettingUse in the videoEffect
25 stepsFaster generationGet a result sooner
40–50 stepsThe author’s preferred higher-quality rangeSpend more sampling time and check whether detail and noise improve
bf16Preferred when VRAM permitsThe author prefers the result at this precision
Lower-precision quantizationReduce resource requirementsCompare visual quality as well as whether the model runs

The video also notes that negative prompts do not have the expected effect with its CFG=1 configuration. When something looks wrong, check the positive prompt, sampling settings and model capability instead of only adding more negative-prompt terms.

3. Editing: distinguish output dimensions from the resolution budget

These two settings are easy to confuse.

Output dimensions determine the shape of the final canvas. For example, to turn a portrait input into a landscape character-view sheet, specify the output width and height instead of inheriting the input aspect ratio. The video demonstrates this switch to a landscape canvas.

The resolution budget affects the image size used for processing. The video shows these values:

BudgetDescription in the video
1024Lower processing resolution
1536The author’s usual compromise
2048Higher resolution and higher resource requirements
0Use the image’s original resolution; check how large the input is

For a large input, using 0 may bring its full memory demands into the workflow. Check the source dimensions before keeping the original resolution. The demonstrated default is 1536.

Start editing with clear, ordinary instructions stating what to preserve and what to change. The video’s editing workflow does not enable prompt rewriting by default; the author finds it sufficient for most demonstrated tasks. Add rewriting when more detailed instructions are useful.

4. What a transparent image means

The goal is an output with an actual alpha channel, not an image that merely depicts a checkerboard. For generation, explicitly request a transparent background. For editing, state which background to remove or change, then inspect the final file’s alpha channel.

The video demonstrates transparent-asset generation and background removal. The official project also describes RGBA generation; see the model project for examples.

What worked well, and what still needs inspection

  • Scenes and lighting: the author likes the compositions, lighting and more three-dimensional-looking images.
  • Posters: the overall layout can look good, but larger amounts of text may still contain garbled characters. Read each part of the generated text.
  • People: fingers and limbs can still be wrong. A pleasing overall appearance does not replace checking anatomy.
  • Editing: character views, comic coloring and style changes are worth trying, but some inputs may not work well.
  • Prompt language: the author finds English more effective in this test. Compare the same editing request in Chinese and English.

For a repeatable editing effect such as anime-to-live-action conversion, continue with Qwen Image 2.1 Editing LoRA: Anime to Live Action.

Resources