Qwen Image 2.1 Consistency LoRA

Update (2026-09-28): I trained this on a bunch of edit pairs to reinforce pixel consistency, and it turns out the dataset might have overfit it on certain types of edits. I've heard it ignores movement, like a new pose or a turned head. I plan to try a v2 once I get the dataset put together. Until then you can try out the two workflows I uploaded:

I made them with my nodes (AusBoss, 2.3.0 or newer). Feel free to swap out my nodes with whatever nodes you want.

Qwen Image 2.1 edits don't stay where you put them. Ask for a watercolor, a comic or an anime version and the picture comes back a few percent taller, sometimes nudged sideways, different for every seed: a waistband 40 px lower, a face that no longer lines up with the original. On a local edit it also repaints things you didn't ask about (hair strands, signs, texture). This LoRA makes the edit on the original's frame: same prompt, same edit, nothing moves.

No trigger word. Load it and write your edit instruction as usual. Also on Civitai.

CRT wall, watercolor: without the LoRA the painting is drawn 5% taller; with it, the waistband stays on the original's line

The dashed line marks where a feature sits in the original; the arrow is how far it moved. Same prompt and seed in every column, rendered in ComfyUI with the Comfy-Org INT8 Qwen Image 2.1 model (25 steps, CFG 1, euler/simple). None of these pictures are in the training set.

Street corner, comic book: plain Qwen stretches the scene 3.9%; the cone tip drops 21 px

Coastal road, comic book: plain Qwen draws the scene 5% taller and the horizon rises 25 px; with the LoRA it stays on the line

Street corner, new cardigan: without the LoRA 9% of the rest of the picture is repainted; with it, 2%

Local edit: both make the change, but without the LoRA Qwen also repaints hair, the storefront sign and the traffic light. The change maps compare each render with the original after lining it up, so they show repainting, not drift.

Files

file notes
qwen-image-2.1-consistency.safetensors step 1500, start here. Keeps Qwen's own look.
qwen-image-2.1-consistency-2000.safetensors step 2000: the tightest alignment; paintings come out a little paler.

Rank 32, ComfyUI key format (diffusion_model.transformer_blocks.*), all 384 tensors load onto the Comfy-Org Qwen Image 2.1 weights.

1500 or 2000

held-out edits, plain Qwen 2.1 vs the LoRA no LoRA 1500 2000
restyles (36): worst corner off, median 24.3 px 1.6 px 0.9 px
restyles: worst corner off, worst case 61.1 px 12.9 px 3.5 px
restyles under 3 px off 3 % 75 % 97 %
restyles (15): look compared with plain Qwen's (colour distance, lower = same look) 7.0 * 6.3 10.9
relights and local edits (12): worst corner off, median 0.7 px 0.1 px 0.1 px

* Two seeds of plain Qwen differ from each other by 7.0, so 1500's restyles look as much like plain Qwen as plain Qwen looks like itself. 2000 lines up more tightly but drifts the style: whiter paper, less colour on ink and gouache.

Only the thing you asked for changes

18 recolor and remove edits on held-out photos of people (a backpack, a hat, a surfboard, a jacket colour). Every render was lined up with its original first, so this counts repainting, not drift:

outside the edited object no LoRA 1500 2000
pixels that changed noticeably (dE > 5) 12.8 % 7.5 % 7.5 %
PSNR against the original 27.7 dB 31.6 dB 31.7 dB

For reference, encoding and decoding a picture through the Qwen 2.1 VAE alone changes about 2 % of the pixels (37 dB), so that is the floor. The edit itself happened in all 18 with and without the LoRA.

ComfyUI

The pack's Qwen Image 2.1 Edit example is the graph these numbers come from. To add the LoRA to any Qwen Image 2.1 edit graph:

  1. LoraLoaderModelOnly right after the model loader, strength 1.0 (lower strengths let some drift back).
  2. Text Encode Qwen Image 2.1: your picture as image_1, resolution 0, and the edit instruction as the prompt.
  3. KSampler on the encoder's latent output: 25 steps, CFG 1, euler / simple, denoise 1. Sample on that latent. A latent of any other size makes Qwen zoom by the size ratio, and no LoRA can undo that.
  4. VAE Decode -> Split Image with Alpha (the Qwen 2.1 VAE decodes RGBA).

Training

  • ostris/ai-toolkit, arch: qwen_image_2, Comfy-Org INT8 convrot base, reference kept at the target's size (match_target_res).
  • 950 edit pairs from 257 pictures: 110 portraits we rendered with Qwen Image 2.1 for this, 133 Unsplash photos (via unsplash-lite) and 14 hand-picked pictures from the author's reference folders.
  • Each pair started as a real Qwen Image 2.1 edit of one of those pictures: restyles, relighting, seasons, photo looks, new backgrounds, outfits, recolors, hair, accessories, objects. The edit's drift was measured and taken out, so every target sits on its source's frame:
    • forward pairs (picture -> edit): the picture is moved onto the edit's frame and the edit is the untouched target;
    • reverse pairs (edit -> picture, "turn this watercolor into a photo"): the edit is moved onto the picture's frame and the untouched original is the target;
    • local edits keep the original's pixels everywhere outside the edited thing. The target itself is never resampled. Restyles were rendered with two seeds and one kept. Every edit was checked by eye and 33 were left out (weak restyles, paintings that bent the picture, subjects who could read as under 18); automatic checks dropped 76 more pairs that were still off after the fix.
  • Captions are the edit instructions themselves, about a third of them followed by a "keep the composition" style phrase. Caption dropout 0.
  • Rank 32 / alpha 32, AdamW8bit, LR 1e-4 constant, batch 1, shift timesteps, resolution: 1408 (pairs of 0.9-1.5 MP on the 32 px grid, so nothing was rescaled), 3000 steps on one H100 (~1.9 s/step), checkpoints every 250. Step 500 already works; 1500 is the sweet spot.

Limits

  • Tested at about 1 MP with the settings above. Not yet tested with 4- or 8-step turbo LoRAs, at 2 MP, or at CFG above 1.
  • Comic and anime restyles redraw every outline, so some shapes still move a few pixels on their own (the worst restyle at 1500 was 12.9 px off at a corner).
  • It keeps the picture in place; it doesn't make weak edits stronger. If plain Qwen won't do an edit, this won't either.
  • The instructions it was trained on are English.

License

A LoRA for Qwen Image 2.1, which is released under the Qwen Research License; use of the base model, and of this LoRA with it, follows that license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for ausboss/Qwen-Image-2.1-Consistency-LoRA

Adapter
(98)
this model

Spaces using ausboss/Qwen-Image-2.1-Consistency-LoRA 4