LayerVerse: Finding the Sweet Spot
for KV-Injection in Training-Free Image Editing

ECCV 2026

1 FusionBrain Lab 2 AXXX 3 Applied AI Institute 4 Lomonosov Moscow State University

Corresponding author: zhirnov.m.d@gmail.com

Source image

FireFlow

FireFlow + LayerVerse

Car → motorcycle

Source image

FireFlow

FireFlow + LayerVerse

Change material to fabric

Source image

FireFlow

FireFlow + LayerVerse

Delete rainbow

LayerVerse jointly selects block identity and injection scope: masked background protection and sparse global structural anchors.

Abstract

For image editing with Multimodal Diffusion Transformers (MM-DiT), training-free methods face a critical trade-off: precise background preservation limits editability, while high prompt fidelity degrades the original scene and object structure. We argue that this tension stems from suboptimal Key-Value (KV) injection strategies. Global KV-injection rigidly over-preserves the source, whereas strictly masked injection acts as localized inpainting, destroying the edited object’s structural identity.

We introduce LayerVerse, an optimization framework that resolves this fundamental dilemma of layer selection for KV-injection by assigning specific layers to two distinct roles: Masked Injection to precisely protect the background, and Global Injection to anchor the edited object’s structure. Formulating optimal layer allocation as a combinatorial problem, we model the editing error based on pairwise layer interactions and solve it globally via Mixed-Integer Linear Programming (MILP). Supported by lightweight, single-step automatic masking, LayerVerse seamlessly integrates into training-free, inversion-based editors, achieving highly competitive results on the PIE-Bench benchmark.

Method

LayerVerse Finds the Right Role for Each Block

Jointly select block identity and injection scope: inactive, masked, or global.

Masked blocks preserve the background; sparse global blocks anchor object structure.

Modeling the editing error

Individual block effects are not independent. LayerVerse models the editing error with a second-order surrogate that includes pairwise interactions.

E(x) is approximately C plus the sum of w_i x_i and the sum of W_ij x_i x_j for i less than j.

wiw_i: individual block effectWi,jW_{i,j}: pairwise synergy or conflict

The coefficients are estimated on a hold-out dataset using no-injection, single-block, and block-pair evaluations.

Optimal role assignment with MILP

Combine the calibrated background and structure objectives under fixed block budgets. Each block is inactive, masked, or global.

Minimize background error E_bg(m+s) plus lambda times structure error E_str(s), with N active blocks and N_str global anchors.

mim_i: masked injectionsis_i: global injection

Mixed-Integer Linear Programming finds the globally optimal assignment for the calibrated surrogate, without exhaustive search.

Auxiliary Automatic Masking for masked KV-injection
Input: source image Z0src, source and target prompts. Parameters: optimal timestep t* and block subset B*. Output: binary editing mask M. 1. Z1 ← Invert(Z0src). 2. Zt* ← (1−t*)Z0src + t*Z1. 3. Procedure ExtractMask(p). 4. Ap ← ForwardPass(Zt*, p). 5. Average cross-attention over B*. 6. Return OtsuBinarize(Smooth(average attention)). 7. End procedure. 8. Msrc ← ExtractMask(psrc). 9. Mtgt ← ExtractMask(ptgt). 10. M ← MorphologicalClose(Msrc ∨ Mtgt). 11. Return M.

One-time architecture calibration. The resulting backbone-specific allocation is reused across editing solvers without recalibration.

Block Allocation and Calibration Details

Optimized block allocation

Backbone / budget Masked KV0-based indices Global KV0-based indices
FLUX.1-[dev]N = 8 · Nstr = 1 Double [1] · Single [32, 33, 34, 35, 36, 37] Double [2]
SD3.5-mediumN = 13 · Nstr = 4 [0, 1, 3, 4, 12, 14, 17, 18, 19] [20, 21, 22, 23]

Calibration dataset

115 held-out MS-COCO images. Error measurements use a standard Euler solver. The same subset is used to select masking parameters offline by maximizing IoU against ground-truth masks.

Computational cost

FLUX.1-[dev] · H100 · one-time offline calibration

Error evaluation · 115 images
∼348 GPU-hours
Error evaluation · 10 images
∼30 GPU-hours
Subsequent MILP solve
∼1.97 seconds

Qualitative Results

Comparisons on PIE-Bench using FLUX.1-[dev] and SD3.5-medium.

Quantitative Evaluation on PIE-Bench

Lowest Average Rank: 2.88 on FLUX.1-[dev]  ·  2.44 on SD3.5-medium

Scroll horizontally to view all metrics.

Quantitative Evaluation on PIE-Bench · FLUX.1-[dev]

Method LV Structure Background Preservation CLIP Similarity User Alignment Rank
Distance ↓ PSNR ↑ LPIPS ↓ MSE ↓ SSIM ↑ Whole ↑ Edited ↑ HPSv3 ↑ EditScore ↑ Avg. ↓
RF-inversion × 42.00 20.15 177.11 141.22 70.64 24.57 21.75 6.36 4.16 5.88
FlowEdit × 27.96 21.92 112.99 95.68 83.35 25.41 22.29 6.81 4.99 4.13
DNAEdit × 17.00 25.18 86.78 48.36 87.19 24.78 21.77 5.80 4.68 4.50
UniEdit-Flow × 8.82 29.74 55.35 18.16 91.12 24.59 21.29 4.27 4.73 3.50
RF-Solver × 11.87 27.81 73.21 28.17 88.35 23.74 20.53 4.93 3.71 5.88
RF-Solver 9.45 30.16 37.62 18.01 92.66 24.30 21.23 5.16 4.93 3.38
FireFlow × 9.48 28.33 64.59 24.51 89.23 23.68 20.45 4.88 3.60 5.88
FireFlow 9.36 30.15 37.82 18.05 92.61 24.39 21.30 5.07 5.05 2.88

LV: LayerVerse. Best results are bold; second-best results are underlined. ↓ / ↑ indicates lower / higher is better.

Quantitative Evaluation on PIE-Bench · SD3.5-medium

Method LV Structure Background Preservation CLIP Similarity User Alignment Rank
Distance ↓ PSNR ↑ LPIPS ↓ MSE ↓ SSIM ↑ Whole ↑ Edited ↑ HPSv3 ↑ EditScore ↑ Avg. ↓
FTEdit × 21.65 23.45 92.85 62.38 85.84 24.47 21.13 4.50 4.33 5.50
FlowEdit × 23.16 23.24 93.58 69.67 85.10 26.59 23.15 6.47 5.92 3.50
FSI-Edit × 14.02 26.60 84.43 36.79 86.31 25.49 22.01 5.66 4.82 3.63
DNAEdit × 13.64 26.65 74.50 32.64 88.91 25.44 22.19 5.41 5.16 2.50
RF-Solver 13.81 26.50 51.46 36.63 89.60 25.14 22.24 4.99 4.91 3.44
FireFlow 13.52 26.58 50.85 35.95 89.70 25.11 22.25 5.01 4.93 2.44

LV: LayerVerse. Best results are bold; second-best results are underlined. ↓ / ↑ indicates lower / higher is better.

Ablations

Both Injection Roles Are Necessary

Global-only KV over-constrains the edit. Masked-only KV preserves the background but loses object structure. LayerVerse combines both roles for the best structure–editability balance.

Scroll sideways to compare all injection policies.

Source image

FireFlow

Only global KV

Only masked KV

FireFlow + LayerVerse

Visual ablation of KV-injection strategies.

Quantitative Comparison of Injection Roles

Injection Type Structure CLIP Similarity User Alignment
Distance ↓ Whole ↑ Edited ↑ HPSv3 ↑ EditScore ↑
Standard FireFlow 9.48 23.68 20.45 4.88 3.59
Only Global KV 11.01 24.28 20.99 4.69 4.23
Only Masked KV 19.03 25.03 22.06 5.64 5.91
LayerVerse (Combined) 9.36 24.39 21.30 5.09 5.05

Pairwise Terms Are Necessary

The quadratic model outperforms the independent linear model in reconstruction quality at every tested budget (N = 2–15). Block synergies and conflicts make joint selection necessary.

Quadratic landscape analysis.

Calibration Data Sensitivity and Computational Cost

The structural anchor and core masked blocks remain stable under resampling. The imputed average-rank analysis shows competitive performance from b ≥ 10 calibration images, reducing offline error-evaluation cost from approximately 348 to 30 H100 GPU-hours.

Data sensitivity analysis: block stability, swap distance, and imputed average rank.

Data Sensitivity on PIE-Bench

Method Structure Background Preservation CLIP Similarity User Alignment Rank
Distance ↓ PSNR ↑ LPIPS ↓ MSE ↓ SSIM ↑ Whole ↑ Edited ↑ HPSv3 ↑ EditScore ↑ Avg. ↓
UniEdit-Flow 8.82 29.74 55.35 18.16 91.12 24.59 21.29 4.27 4.73 3.50
FireFlow + LV 9.21±0.37 30.29±0.21 36.39±2.08 17.64±0.56 92.77±0.19 24.39±0.01 21.28±0.03 5.10±0.03 5.01±0.06 2.88±0.00

± is the maximum metric deviation across the bootstrap candidate pool covering ≥99.9% of empirical mass.

BibTeX

Coming Soon

The publisher-provided BibTeX entry will be added here once available.