~/blog/qwen-image-21-same-face-8-outfits-rtx-5090

One face, eight outfits: keeping a character consistent with Qwen-Image 2.1 on an RTX 5090

❯ cat --toc

TL;DR

One reference photo plus a prompt that says "same woman, only change her clothing" gave me 8 outfits from Qwen-Image 2.1 on an RTX 5090. The face, the mole under her left eye, the wooden hairpin and the background held in all 8. Swapping locations instead, or both at once, works too. An 8-character Traditional Chinese sign came out right, and transparent PNG is built in. Two misses: a landmark hid one letter of a poster headline, and a menu drew 亞 in its simplified form once.

Qwen-Image 2.1 on an RTX 5090: the same woman in eight different outfits, face unchanged

The short version: one face through new outfits and new places, cut stop-motion style.


Before this, keeping one character consistent across images meant training a character LoRA (a small add-on weight file that teaches the model one specific face). Without it, asking for the same person in different clothes usually gave me a different person.

Qwen-Image 2.1 is an open-weight image model from the Qwen team, released on September 20, 2026. The weights are free to download under Qwen's research license; commercial use needs a separate license from Qwen. It accepts reference images, meaning you hand it an existing picture and tell it what to keep. I wanted to know if that could replace the LoRA step. For outfit changes, it can.

Eight outfits, one face: the mole and the hairpin survived every edit

I started from one reference image: a young woman in a pale-blue qipao in an old Taiwanese tea house. Then I asked for eight edits: a pink Song-dynasty Hanfu with embroidered plum blossoms, a Taiwanese high-school uniform, a charcoal business suit, a chef jacket with a tall hat, a NASA-style spacesuit with the helmet under her left arm, a camel winter coat with a red scarf, a pastel idol stage costume, and a black wuxia outfit with a sword on her back.

Three-by-three grid: the reference image top-left and eight outfit edits of the same woman in the same tea house

The background didn't move either. Red lanterns, lattice windows, brick wall and tea table sit in the same place in all eight, and so does she: same spot, same size in the frame. That part took a second attempt, covered in the how-to.

The face is the part that matters, so I cropped all nine.

Three-by-three grid of face close-ups from the reference and all eight outfits, the mole under the left eye visible in each

The mole under her left eye is there in all nine. In the chef image, the wooden hairpin still pokes out beside the hat. I didn't train anything for this.

How to do it: one reference image, one template, a prompt that lists what stays

This section is for running it on your own machine, so expect file names and settings. If you only want to see what it can do, skip to the next section.

Flow diagram: the reference image plus the instruction same woman, only change her clothing go into Qwen-Image 2.1, which outputs eight new outfits with the same face

Step 1: get ComfyUI and three files

ComfyUI is a free, local, node-based image generation app: you wire boxes together into a pipeline, and the official Qwen-Image 2.1 templates come pre-wired. I used ComfyUI 0.37.0. Any version that ships the Qwen-Image 2.1 templates works. The files come from Comfy-Org/Qwen-Image-2.1. The diffusion model is the one to pay attention to:

FolderFileSize
diffusion_models/qwen_image_2.1_int8_convrot.safetensors6.76 GiB
text_encoders/qwen3vl_8b_w4a8.safetensors5.88 GiB
vae/qwen_image_2.1_vae_bf16.safetensors0.63 GiB

The int8 model is Comfy-Org's quantized version. In my own same-prompt comparison its quality was close to bf16, at half the file size, and it fits easily in the 5090's 32 GB. The text encoder reads your prompt and the reference image. I actually ran a community de-refusal ("heretic") variant of the w4a8 encoder, but every prompt here is ordinary and the official one is fine.

Step 2: load a reference image into the edit template

I generated my reference with Qwen-Image 2.1 too: the same description with three seeds (the random number that decides which variation you get), and I picked one of the three. Your own photo works just as well.

Open the ComfyUI template "Qwen Image 2.1 Image Edit". The template preselects a different text encoder (qwen3vl_8b_int8_convrot), so switch its loader to the qwen3vl_8b_w4a8 file you downloaded, and leave the template's optional prompt-rewriting branch off. Then wire the reference into the image_1 input of the "Text Encode Qwen Image 2.1" node, and refer to it as <image1> in the prompt. The official templates wire up to 10 reference images. This post only uses one.

The template has a custom_size switch. Leave it off, which is the default. Off means the sampler starts from the latent output of Text Encode Qwen Image 2.1, which is sized to match the resized reference. On means it uses whatever width and height you type. My first batch did the equivalent of turning it on: I fed a 1536x2048 canvas while the node had resized the reference to 1344x1760, and her position and size drifted from image to image. Details in the deep dive. I set resolution to 1536, so the reference is resized to 1344x1760 and the outputs come out at that size too.

Step 3: say what stays, lock the framing, then say the one change

This is the suit prompt, verbatim:

The same woman from <image1>, with exactly the same face, the same small mole under her left eye, the same hairstyle with the wooden hairpin, the same standing pose, in exactly the same tea house background with the same red lanterns, windows, brick wall, tea table and lighting. Keep exactly the same framing, crop and camera distance as <image1>. Only change her clothing: she now wears a tailored charcoal-grey business suit jacket and trousers with a white blouse. Keep everything else unchanged. Photorealistic.

The "Keep exactly the same framing..." sentence came later. Without it, the suit and the winter coat pulled the camera back into a head-to-toe shot. Leave shoes out of the outfit description too: mention heels or boots and the model wants the feet in frame.

Everything before the clothing sentence is the same across all eight. You don't need to copy my wording.

Step 4: check each result, reroll the odd one out

Most edits land on the first try. When her face sat in the wrong spot, I ran the same prompt with 3 more seeds and kept the one closest to the reference. That happened for 7 of the 23 edits in this post.

Sampling: 40 steps, euler, simple, cfg 1.0. The official template defaults to 25 steps; the official pipeline guidance is about 40 to 50.

The reverse works too: same qipao, eight places, and the lighting follows

Then I flipped it. I kept the outfit and pose fixed, kept the framing line, and changed only the background and lighting, with "Only change the background and lighting: she now stands in ..." in the prompt. Eight places: a Taiwanese night market, a tropical beach, a snowy mountain village, a rainy neon Tokyo street, an old library, a space station with Earth in the window, a cherry-blossom path, and a desert at sunset.

Three-by-three grid: the same woman in the pale-blue qipao placed in eight different locations

The lighting follows the scene instead of being pasted on. Neon reflects off the wet Tokyo pavement. In the desert, her face picks up a warm sunset tint. Same person in every frame, hairpin included.

Outfit and location together: do it in two passes

For "Hanfu under cherry blossoms" I didn't ask for both changes in one prompt. I changed the outfit first, then used that result as the reference for a background swap. The second prompt is shorter because the outfit is already in the reference:

The same woman from <image1>, with exactly the same face, hairstyle, clothing and standing pose. Keep exactly the same framing, crop and camera distance as <image1>. Only change the background and lighting: she now stands in a park path under full-bloom cherry blossom trees. Photorealistic, natural light matching the new scene.

Two-by-four grid: the reference top-left, then the same woman in Hanfu on a cherry-blossom path, a spacesuit in a space station, a coat in a snowy village, an idol costume on a Tokyo street, a wuxia outfit in the desert, a chef jacket at a night market, and a school uniform in a library

Seven combos, one face, one position. Together with the 16 single-change images, that's enough for a stop-motion style quick-cut short.

Text in images: a Chinese sign came out right, with two misses elsewhere

I prompted for a sign reading 巷口阿伯的燒餅油條 ("Uncle's sesame flatbread and fried dough sticks at the alley entrance"). All eight Traditional Chinese characters came out correct, including dense ones like 燒, 餅 and 條.

A street-food stall photo with a shop sign reading 巷口阿伯的燒餅油條, every character correct

Two other tests each had one miss.

An English travel poster: headline "VISIT TAIPEI AT NIGHT", bottom line "Night markets · Hot springs · Bubble tea". Spelling was correct. But Taipei 101 runs through the headline and hides the N in NIGHT; only its right stroke shows.

A chalkboard coffee menu in Chinese and English, with prices: Ethiopia 180, Kenya 200, Alishan 220, Latte 150. All the English and all four prices were right. The miss: in 衣索比亞 (Ethiopia), 亞 is drawn in its simplified form 亚, while the next line, 肯亞 (Kenya), uses the traditional 亞.

Two crops: the poster headline with the N in NIGHT hidden behind Taipei 101, and the menu lines showing 亚 next to 亞

The model makes up any text you don't specify. The stall signs in the night-market images are Chinese-looking strokes that aren't real characters. Put every word you need in the prompt. For dense small text, add it in a layout app afterward.

Transparent PNG straight out of the model, no background removal

Qwen-Image 2.1 can output an alpha channel (the per-pixel transparency layer in a PNG). Wrap your description in two fixed sentences:

This is an RGBA image with transparency. <description> The image has alpha channel and the background is transparent.

My test was a cat barista sticker holding bubble tea. The output is an RGBA PNG. On a checkerboard and on a dark teal background, the edge is clean with no white fringe.

The cat barista sticker shown on a checkerboard and on a dark teal background, with no white fringe at the edges

Speed: about 5.5 s per edit on the 5090, about 11 s on a modded 2080 Ti

Speed isn't why I'd pick this model, so this is short. Each number is from warm runs, three per test. The two machines use different settings, so the rows aren't a head-to-head comparison. The 5090 column is the machine that made the images in this post.

TaskRTX 5090 (int8, 40 steps)Modded RTX 2080 Ti 22GB (turbo int8, 6 steps)
1024x1024 generate~6.4 s (~1.2 s with a 4-step LoRA)~7.5 s
Edit with reference~5.5 s (768x1024)~10.8 s
1536x2048 generate~22.5 snot run

The modded 2080 Ti is my everyday setup at home. Turing has no bf16, so ComfyUI needs to be switched to fp16 to be fast. Details below.

Deep dive: settings, prompts, two ComfyUI 0.37 gotchas, and raw timings

You can skip this section; nothing above depends on it. It's here for anyone who wants to reproduce the results, or hand the whole post to an LLM.

Machine and settings

  • RTX 5090 32 GB, ComfyUI 0.37.0, torch 2.13.0+cu130, comfy-kitchen 0.2.35.
  • Model files as in the table above. The text encoder is the heretic (de-refusal) variant of qwen3vl_8b_w4a8. The official one refuses some prompts, so I use this one daily. Nothing in this post would be refused.
  • Sampling: 40 steps, euler, simple, cfg 1.0. No speed LoRA on any article image.
  • Resolutions: reference 1536x2048; every outfit, location and combo edit at resolution 1536, output 1344x1760; sign 2752x1536; poster 1536x2048; menu and sticker 2048x2048.
  • Sampler start: the latent output of Text Encode Qwen Image 2.1, not an EmptyLatentImage.
  • Seeds: reference from 202001 to 202003, picked 202002.
  • Outfits: Hanfu, uniform, chef, spacesuit, idol and wuxia one run each (303000, 303001, 303003, 303004, 303006, 303007). Suit and coat, after the prompt fix, 3 seeds each; picked 707000 and 707003.
  • Locations: 808000 to 808007 one run each; the library drifted, so 3 more seeds, picked 909132.
  • Combos: coat + snow 809002, idol + Tokyo 809003, uniform + library 809006 on the first try; Hanfu + cherry blossoms, spacesuit + space station, wuxia + desert, chef + night market 3 seeds each, picked 909022, 909033, 909077, 909099.
  • All images regenerated on 2026-10-09.

The base prompt (text-to-image, no reference)

Full-body photograph of a young Chinese woman in her twenties standing in the exact center of the frame, facing the camera, long black hair in a half-up bun with a simple wooden hairpin, oval face, small mole under her left eye. She stands in an old Taiwanese tea house: dark wooden lattice windows, two hanging red paper lanterns, a red brick wall and a low wooden tea table behind her. She wears a pale-blue cotton qipao. Soft afternoon window light from the left, photorealistic, 50mm lens, natural skin texture.

The location prompt (night market)

The same woman from <image1>, with exactly the same face, the same small mole under her left eye, the same hairstyle with the wooden hairpin, the same pale-blue qipao and the same standing pose. Keep exactly the same framing, crop and camera distance as <image1>. Only change the background and lighting: she now stands in a busy Taiwanese night market at night with glowing food stalls. Photorealistic, natural light matching the new scene.

Why the first batch drifted: the canvas didn't match the reference

For the first batch I wired a 1536x2048 EmptyLatentImage into the sampler. At resolution 1536, Text Encode Qwen Image 2.1 resizes the reference to 1344x1760, so the two sizes disagreed. The node's own tooltip in the ComfyUI source says exactly this about its latent output: "Empty latent on the first reference image's size, to match with sampling as any other size shifts the edit." That latent is empty (all zeros). It carries no pixels from the reference; its job is to make the canvas the same size. The template's custom_size switch, when on, takes the same path my first batch did.

I measured the eye midpoint and eye distance in both batches with the YuNet face detector, after resizing every image to 1344x1760:

BatchImagesEye midpoint xEye midpoint yEye distance
First (my own 1536x2048 canvas)17548–634501–69654.5–83.1 (±18.9%)
Final (node's latent output)24611–627645–66173.5–79.6 (±3.9%)

Both include the reference itself (midpoint 618, 650; eye distance 79.2). The worst first-batch outliers were the suit and the coat (eye distance 58.4 and 54.5, pulled back) and the cherry-blossom shot (x drifted to 548).

The suit and coat still went wide on the right canvas

With the canvas fixed, the suit and coat edits still came out as full-body shots. Rerolling didn't help: 3 seeds each on the same prompt, all six landed at eye distance 51.3–59.3, against 79.2 for the reference.

For the second round I changed two things at once: I added Keep exactly the same framing, crop and camera distance as <image1>. and dropped the heels and boots from the outfit descriptions. Again 3 seeds each; all six came out at 77.7–79.5. I didn't test the two changes separately, so the how-to keeps both.

Three panels: the reference; the suit edit when the prompt mentioned heels and had no framing line, pulled back to a full-body shot; the suit edit with the framing line and no shoes, matching the reference position and size

Small pose drift in the outfit set

Face and background are stable across the outfit set. Hands and stance adjust slightly to the clothes. The prompt says "the same standing pose", but the spacesuit requires holding a helmet under the left arm.

ComfyUI 0.37.0 gotcha 1: two required inputs on the text encode node

In 0.37.0, TextEncodeQwenImage21 requires two inputs, negative_prompt and resolution. A hand-built workflow that leaves either out fails through the API with HTTP 400 and this error type:

required_input_missing

Add both inputs to the node and the request goes through.

Gotcha 2: an identical workflow hits the cache in 0.0 s

Submit the exact same workflow twice and ComfyUI returns its cached result: 0.0 s and the same filename. Change the seed when you benchmark, or when you actually want a new image.

Regular outputs carry an alpha channel too

Even without the transparency sentences, outputs are RGBA PNGs. The alpha isn't all 255; the lowest value I saw was 251, far too small a change to see. Convert to RGB before compositing or putting the image on the web.

Raw speed table

Method: warm, three runs each, seed changed every run, timed from submit to image. The first run is slower because it loads the model and the reference image; the tables in the main text use the later runs.

MachineTaskRuns (s)
RTX 50901024x1024, 4-step LoRA2.26 / 1.22 / 1.20
RTX 50901024x1024, 40 steps6.88 / 6.37 / 6.40
RTX 5090768x1024 edit, 40 steps7.4 / 5.5 / 5.5
RTX 50901536x2048 generate / edit, 40 steps~22.5 / ~32.7
RTX 2080 Ti 22GB1024x1024 generate, Viggle 6-step int88.9 / 7.5 / 7.6
RTX 2080 Ti 22GB1024 edit, Viggle 6-step int816.6 / 10.8 / 10.8

Why no article image uses the 4-step LoRA

The 4-step LoRA is the 4step-lora-r64 file from Viggle/Qwen-Image-2.1-viggle-turbo. The LoRA's README says complicated edits, including identity-preserving ones and instructions with several constraints, can still fall short of the base model, with identity drift.

Getting the 2080 Ti fast: fp16, then the 6-step turbo

The 2080 Ti runs the same repo's v0.3-6step-int8_convrot merged checkpoint at 6 steps, cfg 1.0. Turing has no bf16, so ComfyUI falls back to fp32, which is slow. I added fp16 to ComfyUI's QwenImage21.supported_inference_dtypes and moved to torch 2.9.1+cu130. The same int8 40-step job went from 159 s to 42 s. Then I switched to the 6-step turbo checkpoint, which gives the 7.5 s per 1024x1024 image above.

FAQ

How do I keep the same face when changing a character's clothes in Qwen-Image 2.1?
Feed one reference image into the Text Encode Qwen Image 2.1 node, sample on that node's own latent output so the canvas matches the reference size, and write the prompt in three parts: everything that must stay (face, mole, hairstyle and hairpin, pose, background), the line 'Keep exactly the same framing, crop and camera distance as <image1>.', then 'Only change her clothing' plus the new outfit. That gave me 8 outfits where the face, the mole and the wooden hairpin never changed, and the face barely moved in the frame. No character LoRA needed.
Can Qwen-Image 2.1 render Chinese text correctly?
Mostly. A shop sign reading 巷口阿伯的燒餅油條 came out with all 8 Traditional Chinese characters correct, including dense ones like 燒, 餅 and 條. On a chalkboard menu it drew 亞 in its simplified form 亚 on one line and correctly on the next. Small background text you didn't ask for gets invented, so put every word you need in the prompt.
How fast is Qwen-Image 2.1 on an RTX 5090?
With the int8 model at 40 steps, warm: about 6.4 s for a 1024x1024 image, about 5.5 s for a 768x1024 edit with a reference image, and about 1.2 s at 1024x1024 with a 4-step LoRA. A 1536x2048 text-to-image render took about 22.5 s.
Can Qwen-Image 2.1 make transparent PNGs?
Yes. Wrap the description in two fixed sentences: 'This is an RGBA image with transparency.' before it and 'The image has alpha channel and the background is transparent.' after it. The output is an RGBA PNG with a clean edge and no white fringe on a dark background, so you skip the background-removal step.

Read next

Don't miss the next one

Subscribe, and you won't.

One-click unsubscribe anytime.