提出无需训练的图像反演方法,用极小存储实现高质量重建。
Compression Asymmetry and Trajectory Binding in Noise-Anchored Diffusion Inversion

- 用单个int8噪声锚点+动态权重调度,实现高效反演。
- 相比基线节省400倍存储,PSNR提升3.24 dB。
- 适用于SD1.5和SDXL,支持直接编辑插件使用。
真实图像的扩散反演受质量与成本(计算、存储或逐图优化)之间的紧密权衡制约。本文通过前向高斯噪声锚点定义的扩散轨迹,揭示了有效存储噪声反演的两个机制:第一,扩散噪声存在元素级压缩不对称性——完整维度的int8锚点可保持重建质量,而低维子空间摘要则可靠性差,常在相当或更小负载下崩溃;该元素级对子空间的排序在五种存储噪声反演方法中均成立。第二,反演依赖于轨迹绑定与得分先验耦合:匹配的前向锚点与训练好的得分网络缺一不可,反对纯代数恒等解释。这些发现明确了应存储的内容及其使用方式。由此提出无训练反演原语NARC,仅需存储一个int8潜在锚点,并以固定、噪声水平依赖的锚定权重调度重用:逆向轨迹噪声主导时强锚定,图像细节显现后逐渐放松。在PIE-Bench++上,使用Stable Diffusion 1.5时,NARC超越五个现代非精确基线,且相比PnP DirectInv提升PSNR 3.24 dB,同时存储量仅为后者的约1/400。压缩不对称性、锚点特异性及编辑插件能力亦可迁移至SDXL 1024^2。
原文摘要 · Abstract (English)
Real-image diffusion inversion is governed by a tight quality-cost trade-off, with costs incurred in computation, storage, or per-image optimization. We study this trade-off through the forward Gaussian noise anchor that defines a diffusion trajectory and isolate two mechanisms behind effective stored-noise inversion. First, diffusion noise exhibits an element-wise compression asymmetry: int8 full-dimensional anchors preserve reconstruction, whereas low-dimensional subspace summaries are much less reliable, often collapsing even at comparable or smaller payloads; the element-wise over subspace ordering persists across five stored-noise inversion methods. Second, inversion is trajectory-bound and score-prior coupled: the matched forward anchor and a trained score network are both necessary, arguing against a purely algebraic-identity explanation. Together, these findings specify what to store and how to use it. They lead to Noise-Anchored Reverse Correction (NARC), a training-free inversion primitive that stores a single int8 latent anchor and reuses it with a fixed, noise-level-dependent anchor-weight schedule: strong anchoring when the reverse trajectory is noise-dominated, then relaxed anchoring as image detail emerges. On PIE-Bench++ with Stable Diffusion 1.5, NARC outperforms five modern non-exact baselines and improves PSNR by +3.24 dB over PnP DirectInv while using about 400x less inversion storage than PnP DirectInv. The compression asymmetry, anchor specificity, and editing plug-in also transfer to SDXL 1024^2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。