qwen-image-edit · diff
git:20260831.2ba287a to git:20260831.22f25d6
75 added, 5 removed. Audit A to A.
---
name: qwen-image-edit
description: Build Qwen Image Edit workflows covering model loading, conditioning, LoRAs, prompt patterns, and XY plot testing
globs:
- "**/*.json"
---
# Qwen Image Edit Workflows
## Overview
Qwen Image Edit uses a vision-language model (Qwen2.5-VL) to edit images based on natural language instructions. The model "sees" the source image through CLIP conditioning and generates an edited version.
## Models
### Required Components
| Component | Node | Model Name | Notes |
|-----------|------|------------|-------|
| **UNET** | `UNETLoader` | `qwen_image_edit_2511_bf16.safetensors` | Official 2511 edit model (bf16) |
| **CLIP** | `CLIPLoader` (type=`qwen_image`) | `qwen_2.5_vl_7b_fp8_scaled.safetensors` | Shared across all Qwen models |
| **VAE** | `VAELoader` | `qwen_image_vae.safetensors` | Qwen-specific VAE |
### Alternative UNET Models
| Model | Path | Focus |
|-------|------|-------|
| `qwenImageEditRemix_v10` | `qwenImageEditRemix_v10.safetensors` | Community remix, general editing |
| `qwenUltimateRealism_v11` | `Qwen/imageized/qwenUltimateRealism_v11.safetensors` | Product photography, hyper-realistic |
| `copaxTimeless` | `Qwen/realistic/copaxTimeless_qwenUltraRealistic.safetensors` | Ultra-realistic portraits |
| `qwnImageEdit_v16Bf16` | `Qwen/abliterated/qwnImageEdit_v16Bf16.safetensors` | Abliterated (uncensored) |
## Conditioning Nodes
### TextEncodeQwenImageEditPlusAdvance_lrzjason (Recommended)
From the `qweneditutils` custom node pack. The **Advanced** variant is preferred because it:
- Outputs a **LATENT directly** (no need for separate EmptyLatentImage)
- Has separate **VL-resize** and **non-resize** image slots for fine control
- Supports **target_size** control for output resolution
- Includes a **pad/center/disabled crop** method with pad_info output
```
Required Inputs:
- clip: CLIP
- prompt: STRING — natural language edit instruction
Optional Inputs:
- vae: VAE — needed for image encoding and latent output
- vl_resize_image1-3: IMAGE — images that get VL-resized (downscaled for vision encoder)
- not_resize_image1-3: IMAGE — images kept at full resolution
- target_size: [1024, 1344, 1536, 2048, 768, 512] (default 1024)
- target_vl_size: [392, 384] (default 384)
- upscale_method: [lanczos, bicubic, area]
- crop_method: [pad, center, disabled]
- instruction: STRING — system instruction template (has sensible default)
Outputs (10):
[0] conditioning_with_full_ref: CONDITIONING — use as positive conditioning
[1] latent: LATENT — auto-scaled latent, feed directly to KSampler
[2] target_image1: IMAGE — processed target-size image
[3] target_image2: IMAGE
[4] target_image3: IMAGE
[5] vl_resized_image1: IMAGE — VL-resized version
[6] vl_resized_image2: IMAGE
[7] vl_resized_image3: IMAGE
[8] conditioning_with_first_ref: CONDITIONING — conditioning with only first ref
[9] pad_info: ANY — padding info for later unpadding
```
**Key advantage**: Output [1] (latent) eliminates the need for a separate `EmptyLatentImage` or `VAEEncode` node. The Advanced node handles latent creation internally at the correct resolution.
### Other Conditioning Variants
- **TextEncodeQwenImageEditPlus** (Phr00t v2, built-in) is simpler: 4 image inputs, outputs only CONDITIONING. Requires separate EmptyLatentImage. Good for quick edits.
- **TextEncodeQwenImageEditPlus_lrzjason**: 5 image inputs, resize toggles, but less control than Advance
- **TextEncodeQwenImageEditPlusPro_lrzjason**: Per-image VL resize selection via `vl_resize_indexs` string, `main_image_index` control
## Lightning LoRAs (Fast Generation)
### 4-Step Lightning (2511 Edit)
```json
{
"class_type": "LoraLoaderModelOnly",
"inputs": {
"model": ["<unet_node>", 0],
"lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors",
"strength_model": 1.0
}
}
```
**Settings**: steps=4, cfg=1.0, sampler=euler, scheduler=simple, denoise=1.0
### 4-Step Lightning (General Qwen)
For non-edit models (txt2img, 2512):
- `Qwen-Image-Lightning-4steps-V1.0.safetensors` (strength 1.0)
### 8-Step Lightning
- `Qwen-Image-Lightning-8steps-V1.0.safetensors`, higher detail than 4-step
## Sampler Settings
| Preset | Steps | CFG | Sampler | Scheduler | Denoise | LoRA |
|--------|-------|-----|---------|-----------|---------|------|
| Lightning 4-step (2511 edit) | 4 | 1.0 | euler | simple | 1.0 | 2511-Lightning-4steps |
| Lightning 8-step | 8 | 1.0 | euler | simple | 1.0 | Lightning-8steps |
| Standard edit | 40 | 4.0 | euler | simple | 0.75 | none |
| Quality edit | 50 | 4.0 | euler | simple | 0.5-0.8 | none |
> **The sub-1.0 denoise rows REQUIRE a `VAEEncode` latent.** A denoise low enough to
> shorten the sampling schedule — which 0.5-0.8 certainly is — keeps part of the
> incoming latent, so that latent has to BE the source image. Wire `latent_image` from
> a `VAEEncode` of the source (or from a node that emits a source-derived latent, like
> `TextEncodeQwenImageEditPlusAdvance_lrzjason` output [1]). Pairing these rows with an
> `EmptyLatentImage` runs clean and returns a flat, near-uniform field — an empty latent
> has no source content to preserve. Feeding the reference through
> `TextEncodeQwenImageEditPlus` does **not** rescue it: that image rides on CONDITIONING,
> which steers denoising but never seeds the sampler's starting state.
**Denoise for editing**: Lower denoise = closer to source — *provided the latent IS the source*. 0.5-0.8 range for standard editing on a `VAEEncode` latent. Lightning uses 1.0 (model handles fidelity internally).
## Resolutions
- Qwen operates at ~1.6 megapixels natively:
+ > **This table is for Qwen-Image TEXT-TO-IMAGE. Do not pick an edit graph's output
+ > size from it.** An edit graph's geometry is decided by the SOURCE image, not by you
+ > — see "Resolution on an edit graph" below. Choosing 1104x1472 here for an edit was
+ > #2681.
+ Qwen-Image operates at ~1.6 megapixels natively:
+
| Aspect | Resolution | Use Case |
|--------|-----------|----------|
| Square | 1328x1328 | General |
| Portrait 3:4 | 1104x1472 | Portraits |
| Portrait 9:16 | 928x1664 | Phone format |
| Landscape 4:3 | 1472x1104 | Landscape scenes |
| Landscape 16:9 | 1664x928 | Widescreen |
| Video-ready | 832x480 | For WAN 2.2 FLF pipeline |
**For video pipelines**: Use 832x480 to match WAN 2.2's default resolution.
+ ### Resolution on an edit graph
+
+ `TextEncodeQwenImageEdit` and `TextEncodeQwenImageEditPlus` do not take a size. They
+ scale every reference image to a hard-coded `int(1024 * 1024)` px — **~1.05 MP, at the
+ source's own aspect ratio** — VAE-encode it, and hand it to the model as a reference
+ latent (`comfy_extras/nodes_qwen.py`).
+
+ The model then lays the reference tokens and the tokens it is generating on **one
+ shared, centred coordinate grid** (`comfy/ldm/qwen_image/model.py`, `process_img`), so
+ reference position (i, j) and output position (i, j) mean the same place only when the
+ two grids are the same size. That is the whole reason the official templates run the one
+ image through `FluxKontextImageScale` and then feed the sampler a `VAEEncode` of *that*
+ scaled image — both branches then see the same pixels at the same ~1 MP scale and the
+ same aspect. Every `PREFERRED_KONTEXT_RESOLUTIONS` entry is ~1.05 MP for the same reason.
+
+ It is agreement to within the encoder's round-to-8, not exact equality, and the
+ difference is worth knowing precisely. `FluxKontextImageScale` snaps to a preferred pair;
+ the encoder then renormalises *that* to its own 1,048,576 px budget. For 12 of the 18
+ preferred pairs the two land on the same latent grid. For the other 6 — 688x1504,
+ 800x1328, 832x1248 and their landscape mirrors — the reference lands one latent row or
+ column off: at 800x1328 the sampler's grid is 100x166 and the reference's is 99x165.
+ ComfyUI's own bundled 2511 template does exactly this, so a sub-patch offset is evidently
+ fine in practice. **The failure this page is about is one of SCALE, not of rounding** — an
+ empty latent at 1104x1472 sits 1.24x away linearly, not one row.
+
+ So on an edit graph you do not choose a resolution — you inherit one:
+
+ - **Right:** `LoadImage` -> `FluxKontextImageScale` -> (`TextEncodeQwenImageEditPlus`
+ *and* `VAEEncode`) -> KSampler `latent_image`.
+ - **Wrong:** an `EmptyLatentImage` at a size from the table above. Its dimensions are
+ literals; the reference's are computed from the source when the graph runs. At
+ 1104x1472 (1.63 MP) against a 1.05 MP reference the grids are 1.24x apart linearly,
+ the model cannot copy detail across them, and it re-synthesises the subject instead —
+ materials come back looking plastic/CGI and printed detail comes back as a generic
+ shape (#2681). Nothing errors; the image just is not the edit you asked for.
+
+ `create_workflow (action:"validate")` now flags this pairing as
+ `edit_reference_empty_latent`.
+
## Prompt Patterns
### Edit Instructions (Natural Language)
```
"Change the black cat into a cute girl with a black bodysuit and jeans"
"Make the sky a dramatic sunset with orange and purple clouds"
"Add a red sports car parked in front of the house"
"Remove the person on the left and fill with the background"
```
### Multi-Angle LoRA (qwen-image-edit-2511-multiple-angles-lora)
Uses `<sks>` token with structured angle/distance prompts:
```
<sks> front view eye-level shot close-up
<sks> front-right quarter view low-angle shot medium shot
<sks> back view elevated shot wide shot
```
**Template**: `<sks> {direction} view {angle} shot {distance}`
Directions: front, front-right quarter, right side, back-right quarter, back, back-left quarter, left side, front-left quarter
Angles: low-angle, eye-level, elevated, high-angle
Distances: close-up, medium shot, wide shot
## Negative Conditioning
Always use `ConditioningZeroOut` for negative conditioning with Qwen edit:
```json
{
"class_type": "ConditioningZeroOut",
"inputs": { "conditioning": ["<positive_cond_node>", 0] }
}
```
## Complete Workflow: Lightning Edit (Advanced Node)
Uses `TextEncodeQwenImageEditPlusAdvance_lrzjason`, which outputs the latent directly, so no EmptyLatentImage is needed.
```json
{
"1": { "class_type": "UNETLoader", "inputs": { "unet_name": "qwen_image_edit_2511_bf16.safetensors", "weight_dtype": "default" }},
"2": { "class_type": "LoraLoaderModelOnly", "inputs": { "model": ["1", 0], "lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors", "strength_model": 1 }},
"3": { "class_type": "CLIPLoader", "inputs": { "clip_name": "qwen_2.5_vl_7b_fp8_scaled.safetensors", "type": "qwen_image" }},
"4": { "class_type": "VAELoader", "inputs": { "vae_name": "qwen_image_vae.safetensors" }},
"5": { "class_type": "LoadImage", "inputs": { "image": "<source_image.png>" }},
"6": { "class_type": "TextEncodeQwenImageEditPlusAdvance_lrzjason", "inputs": {
"clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0],
"vl_resize_image1": ["5", 0],
"target_size": 1024, "target_vl_size": 384,
"upscale_method": "lanczos", "crop_method": "pad"
}},
"7": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["6", 0] }},
"8": { "class_type": "KSampler", "inputs": {
"model": ["2", 0],
"positive": ["6", 0],
"negative": ["7", 0],
"latent_image": ["6", 1],
"seed": 42, "steps": 4, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
}},
"9": { "class_type": "VAEDecode", "inputs": { "samples": ["8", 0], "vae": ["4", 0] }},
"10": { "class_type": "SaveImage", "inputs": { "images": ["9", 0], "filename_prefix": "qwen_edit" }}
}
```
**Key connections**:
- `"latent_image": ["6", 1]`: KSampler gets its latent directly from the Advanced node's output [1]
- `"positive": ["6", 0]`: conditioning_with_full_ref from output [0]
- `"vl_resize_image1": ["5", 0]`: source image goes into VL-resize slot (downscaled for vision encoder)
### Simpler Alternative (Phr00t v2)
If `qweneditutils` custom node is unavailable, use the built-in `TextEncodeQwenImageEditPlus` with a separate `EmptyLatentImage`:
```json
{
"6": { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
"clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5", 0]
}},
"8": { "class_type": "EmptyLatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size": 1 }}
}
```
Replace node 6 and add node 8. KSampler latent_image connects to `["8", 0]` instead of `["6", 1]`.
- **Set `denoise` to 1.0 on this path.** `TextEncodeQwenImageEditPlus` emits CONDITIONING only, so the `EmptyLatentImage` here is genuinely empty and carries no source pixels. A sub-1.0 denoise over it produces a flat texture instead of the edited image (#2678). If you want a sub-1.0 denoise, drop the `EmptyLatentImage` and feed `latent_image` from a `VAEEncode` of the source image instead:
+ **Do not leave that `EmptyLatentImage` wired to the sampler.** It is shown above only
+ because it is what the Plus encoder's own signature leaves you needing, and it fails in a
+ different way on each side of denoise 1.0:
+ - Below 1.0 it renders a flat, near-uniform field, always. The encoder emits CONDITIONING
+ only, so the latent is genuinely empty, and a truncated sigma schedule exists to
+ preserve an incoming latent that here has nothing in it (#2678).
+ - At 1.0 it renders a plausible image that is not your source — *unless* its width and
+ height happen to equal the geometry the encoder derived, which is ~1.05 MP at the
+ source's aspect ratio. A 1024x1024 empty latent over an exactly-square source does line
+ up, and is fine. 1104x1472 over that same source does not: the model aligns reference
+ and output on one shared grid, cannot copy across grids that far apart, and
+ re-synthesises instead (#2681). See "Resolution on an edit graph".
+
+ The second case is the trap, because it depends on a source you may not have looked at
+ and it fails silently. A `VAEEncode` removes the coincidence — it cannot be the wrong
+ size, because it is derived from the same pixels the encoder saw.
+
+ Feed `latent_image` from a `VAEEncode` of the same image you gave the encoder, scaled
+ once up front so both branches see the same pixels at the same scale:
+
```json
{
- "8": { "class_type": "VAEEncode", "inputs": { "pixels": ["5", 0], "vae": ["4", 0] }}
+ "5b": { "class_type": "FluxKontextImageScale", "inputs": { "image": ["5", 0] }},
+ "6": { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
+ "clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5b", 0]
+ }},
+ "8": { "class_type": "VAEEncode", "inputs": { "pixels": ["5b", 0], "vae": ["4", 0] }}
}
```
+ KSampler `latent_image` connects to `["8", 0]`. Any denoise is then meaningful: 1.0 for
+ a full edit, sub-1.0 to stay closer to the source.
+
### Basic Variant (Official ComfyUI Example)
The official "Qwen 2511 Edit Simple" example uses newer built-in nodes for model patching and image scaling:
**Additional nodes in the official pipeline:**
- **`ModelSamplingAuraFlow`** (shift=3.1): Flow matching shift applied to the UNET. Used instead of `ModelSamplingSD3`.
- **`CFGNorm`** (strength=1): Normalizes CFG guidance for more stable generation. Applied after `ModelSamplingAuraFlow`.
- **`FluxKontextImageScale`**: Auto-scales input images to the correct resolution for Qwen. No manual size parameters needed.
- **`FluxKontextMultiReferenceLatentMethod`** (method=`index_timestep_zero`): Applied to both positive and negative conditioning. Handles multi-reference latent indexing.
- **`VAEEncode`**: Encodes the scaled image to latent (instead of `EmptyLatentImage`).
**Official pipeline flow:**
```
UNETLoader → [LoraLoaderModelOnly] → ModelSamplingAuraFlow (shift=3.1) → CFGNorm (strength=1) → MODEL
CLIPLoader (qwen_image) → CLIP
VAELoader → VAE
LoadImage → FluxKontextImageScale → scaled_image
├─ TextEncodeQwenImageEditPlus (positive) → FluxKontextMultiReferenceLatentMethod → positive CONDITIONING
├─ TextEncodeQwenImageEditPlus (negative, empty) → FluxKontextMultiReferenceLatentMethod → negative CONDITIONING
└─ VAEEncode → LATENT
KSampler → VAEDecode → SaveImage
```
**Official sampler settings:**
| Variant | Steps | CFG | Sampler | Scheduler | Denoise | LoRA |
|---------|-------|-----|---------|-----------|---------|------|
| Standard | 40 | 4.0 | euler | simple | 1.0 | none |
| Lightning | 4 | 1.0 | euler | simple | 1.0 | 2511-Lightning-4steps |
**Note**: The `FluxKontextMultiReferenceLatentMethod` and `FluxKontextImageScale` nodes may not be needed when using Comfy's official model files directly, but may be required with community-repackaged models.
## XY Plot Technique (from Widgets.json)
For batch-testing multiple edit variations, use the Easy Nodes XY Plot system:
1. **Text Multiline** nodes define parameter lists (e.g., directions, angles)
2. **Split String** breaks them into indexed options
3. **easy textIndexSwitch** selects one at a time
4. **easy promptReplace** substitutes `{X}`, `{Y}`, `{Z}` placeholders in the base prompt
5. **easy XYPlotAdvanced** + **easy XYInputs: PromptSR** drives the sweep
6. **easy pipeIn** bundles model/clip/vae/latent into a pipeline
This produces a grid image showing all combinations, useful for finding the best angle/distance/style for a given subject.
## VRAM Considerations
- Qwen 2511 edit bf16: ~10GB VRAM
- CLIP (fp8): ~7GB VRAM
- VAE: ~200MB
- Total: ~17-18GB, fits comfortably on 24GB GPUs
- **Always `clear_vram` before loading** if switching from another model family
## Tips
1. **Upload source images first** with `upload_image (action:"image")` before building the workflow
- 2. **Match output resolution** to the next pipeline step (e.g., 832x480 for WAN FLF)
+ 2. **Match output resolution** to the next pipeline step (e.g., 832x480 for WAN FLF) — but only on a *generation* graph. On an edit graph the size is the source's; resize the RESULT afterwards instead of sampling at the size you want
3. **Lightning LoRA + denoise 1.0** works well. The model handles structure preservation through conditioning
- 4. For **img2img editing** (denoise < 1.0), use `VAEEncode` on the source image instead of `EmptyLatentImage` — a sub-1.0 denoise over an empty latent decodes to a flat, near-uniform field with no error (#2678). `create_workflow (action:"validate")` now flags this pairing
+ 4. **Take an edit graph's `latent_image` from a `VAEEncode` of the source, not from an `EmptyLatentImage`** — at sub-1.0 denoise the empty latent decodes to a flat, near-uniform field (#2678), and at denoise 1.0 it is right only if its literal size happens to equal the geometry the encoder derived from the source, which is exactly the coincidence a `VAEEncode` removes (#2681). Neither failure errors. `create_workflow (action:"validate")` flags both pairings (`partial_denoise_empty_latent`, `edit_reference_empty_latent`)
5. The **lrzjason Pro variant** is best for multi-image compositions where you need fine control over which images get VL-resized
6. **Use `get_workflow (action:"analyze")`** to understand any saved Qwen edit workflow before modifying or executing it. It returns a structured summary, not raw JSON. Only use `get_workflow` when you need the actual JSON for `enqueue_workflow` or `create_workflow (action:"modify")`.
## Sources
- **Official:** none found.
- **Empirical:** sampler values, wiring, and prompt notes from working graphs in `packs/` and observed renders; not a vendor prompting guide.