Abstract
For image editing with Multimodal Diffusion Transformers (MM-DiT), training-free methods face a critical trade-off: precise background preservation limits editability, while high prompt fidelity degrades the original scene and object structure. We argue that this tension stems from suboptimal Key-Value (KV) injection strategies. Global KV-injection rigidly over-preserves the source, whereas strictly masked injection acts as localized inpainting, destroying the edited object’s structural identity.
We introduce LayerVerse, an optimization framework that resolves this fundamental dilemma of layer selection for KV-injection by assigning specific layers to two distinct roles: Masked Injection to precisely protect the background, and Global Injection to anchor the edited object’s structure. Formulating optimal layer allocation as a combinatorial problem, we model the editing error based on pairwise layer interactions and solve it globally via Mixed-Integer Linear Programming (MILP). Supported by lightweight, single-step automatic masking, LayerVerse seamlessly integrates into training-free, inversion-based editors, achieving highly competitive results on the PIE-Bench benchmark.
Method
LayerVerse Finds the Right Role for Each Block
Jointly select block identity and injection scope: inactive, masked, or global.
Modeling the editing error
Individual block effects are not independent. LayerVerse models the editing error with a second-order surrogate that includes pairwise interactions.
: individual block effect: pairwise synergy or conflict
The coefficients are estimated on a hold-out dataset using no-injection, single-block, and block-pair evaluations.
Optimal role assignment with MILP
Combine the calibrated background and structure objectives under fixed block budgets. Each block is inactive, masked, or global.
: masked injection: global injection
Mixed-Integer Linear Programming finds the globally optimal assignment for the calibrated surrogate, without exhaustive search.
Auxiliary Automatic Masking for masked KV-injection
One-time architecture calibration. The resulting backbone-specific allocation is reused across editing solvers without recalibration.
Block Allocation and Calibration Details
Optimized block allocation
| Backbone / budget | Masked KV0-based indices | Global KV0-based indices |
|---|---|---|
| FLUX.1-[dev]N = 8 · Nstr = 1 | Double [1] · Single [32, 33, 34, 35, 36, 37] |
Double [2] |
| SD3.5-mediumN = 13 · Nstr = 4 | [0, 1, 3, 4, 12, 14, 17, 18, 19] |
[20, 21, 22, 23] |
Calibration dataset
115 held-out MS-COCO images. Error measurements use a standard Euler solver. The same subset is used to select masking parameters offline by maximizing IoU against ground-truth masks.
Computational cost
FLUX.1-[dev] · H100 · one-time offline calibration
- Error evaluation · 115 images
- ∼348 GPU-hours
- Error evaluation · 10 images
- ∼30 GPU-hours
- Subsequent MILP solve
- ∼1.97 seconds
Qualitative Results
Comparisons on PIE-Bench using FLUX.1-[dev] and SD3.5-medium.
Scroll sideways to see all three images, or switch to Before / after.
Quantitative Evaluation on PIE-Bench
Lowest Average Rank: 2.88 on FLUX.1-[dev] · 2.44 on SD3.5-medium
Scroll horizontally to view all metrics.
Quantitative Evaluation on PIE-Bench · FLUX.1-[dev]
| Method | LV | Structure | Background Preservation | CLIP Similarity | User Alignment | Rank | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Distance ↓ | PSNR ↑ | LPIPS ↓ | MSE ↓ | SSIM ↑ | Whole ↑ | Edited ↑ | HPSv3 ↑ | EditScore ↑ | Avg. ↓ | ||
| RF-inversion | × | 42.00 | 20.15 | 177.11 | 141.22 | 70.64 | 24.57 | 21.75 | 6.36 | 4.16 | 5.88 |
| FlowEdit | × | 27.96 | 21.92 | 112.99 | 95.68 | 83.35 | 25.41 | 22.29 | 6.81 | 4.99 | 4.13 |
| DNAEdit | × | 17.00 | 25.18 | 86.78 | 48.36 | 87.19 | 24.78 | 21.77 | 5.80 | 4.68 | 4.50 |
| UniEdit-Flow | × | 8.82 | 29.74 | 55.35 | 18.16 | 91.12 | 24.59 | 21.29 | 4.27 | 4.73 | 3.50 |
| RF-Solver | × | 11.87 | 27.81 | 73.21 | 28.17 | 88.35 | 23.74 | 20.53 | 4.93 | 3.71 | 5.88 |
| RF-Solver | ✓ | 9.45 | 30.16 | 37.62 | 18.01 | 92.66 | 24.30 | 21.23 | 5.16 | 4.93 | 3.38 |
| FireFlow | × | 9.48 | 28.33 | 64.59 | 24.51 | 89.23 | 23.68 | 20.45 | 4.88 | 3.60 | 5.88 |
| FireFlow | ✓ | 9.36 | 30.15 | 37.82 | 18.05 | 92.61 | 24.39 | 21.30 | 5.07 | 5.05 | 2.88 |
LV: LayerVerse. Best results are bold; second-best results are underlined. ↓ / ↑ indicates lower / higher is better.
Quantitative Evaluation on PIE-Bench · SD3.5-medium
| Method | LV | Structure | Background Preservation | CLIP Similarity | User Alignment | Rank | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Distance ↓ | PSNR ↑ | LPIPS ↓ | MSE ↓ | SSIM ↑ | Whole ↑ | Edited ↑ | HPSv3 ↑ | EditScore ↑ | Avg. ↓ | ||
| FTEdit | × | 21.65 | 23.45 | 92.85 | 62.38 | 85.84 | 24.47 | 21.13 | 4.50 | 4.33 | 5.50 |
| FlowEdit | × | 23.16 | 23.24 | 93.58 | 69.67 | 85.10 | 26.59 | 23.15 | 6.47 | 5.92 | 3.50 |
| FSI-Edit | × | 14.02 | 26.60 | 84.43 | 36.79 | 86.31 | 25.49 | 22.01 | 5.66 | 4.82 | 3.63 |
| DNAEdit | × | 13.64 | 26.65 | 74.50 | 32.64 | 88.91 | 25.44 | 22.19 | 5.41 | 5.16 | 2.50 |
| RF-Solver | ✓ | 13.81 | 26.50 | 51.46 | 36.63 | 89.60 | 25.14 | 22.24 | 4.99 | 4.91 | 3.44 |
| FireFlow | ✓ | 13.52 | 26.58 | 50.85 | 35.95 | 89.70 | 25.11 | 22.25 | 5.01 | 4.93 | 2.44 |
LV: LayerVerse. Best results are bold; second-best results are underlined. ↓ / ↑ indicates lower / higher is better.
Ablations
Both Injection Roles Are Necessary
Global-only KV over-constrains the edit. Masked-only KV preserves the background but loses object structure. LayerVerse combines both roles for the best structure–editability balance.
Scroll sideways to compare all injection policies.
Source image
FireFlow
Only global KV
Only masked KV
FireFlow + LayerVerse
Visual ablation of KV-injection strategies.
Quantitative Comparison of Injection Roles
| Injection Type | Structure | CLIP Similarity | User Alignment | ||
|---|---|---|---|---|---|
| Distance ↓ | Whole ↑ | Edited ↑ | HPSv3 ↑ | EditScore ↑ | |
| Standard FireFlow | 9.48 | 23.68 | 20.45 | 4.88 | 3.59 |
| Only Global KV | 11.01 | 24.28 | 20.99 | 4.69 | 4.23 |
| Only Masked KV | 19.03 | 25.03 | 22.06 | 5.64 | 5.91 |
| LayerVerse (Combined) | 9.36 | 24.39 | 21.30 | 5.09 | 5.05 |
Pairwise Terms Are Necessary
The quadratic model outperforms the independent linear model in reconstruction quality at every tested budget (N = 2–15). Block synergies and conflicts make joint selection necessary.
Calibration Data Sensitivity and Computational Cost
The structural anchor and core masked blocks remain stable under resampling. The imputed average-rank analysis shows competitive performance from b ≥ 10 calibration images, reducing offline error-evaluation cost from approximately 348 to 30 H100 GPU-hours.
Data Sensitivity on PIE-Bench
| Method | Structure | Background Preservation | CLIP Similarity | User Alignment | Rank | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Distance ↓ | PSNR ↑ | LPIPS ↓ | MSE ↓ | SSIM ↑ | Whole ↑ | Edited ↑ | HPSv3 ↑ | EditScore ↑ | Avg. ↓ | |
| UniEdit-Flow | 8.82 | 29.74 | 55.35 | 18.16 | 91.12 | 24.59 | 21.29 | 4.27 | 4.73 | 3.50 |
| FireFlow + LV | 9.21±0.37 | 30.29±0.21 | 36.39±2.08 | 17.64±0.56 | 92.77±0.19 | 24.39±0.01 | 21.28±0.03 | 5.10±0.03 | 5.01±0.06 | 2.88±0.00 |
± is the maximum metric deviation across the bootstrap candidate pool covering ≥99.9% of empirical mass.
BibTeX
Coming Soon
The publisher-provided BibTeX entry will be added here once available.