arXiv:2606.24000cs.LGcond-mat.dis-nn2026-06被引 1

通过循环去噪暴露扩散模型中长期记忆的图像,无需额外信息。

Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models

论文配图:Cyclic Denoising Reveals Ultrastable Memories in Diffusion Models
图 1 · 摘自论文原文
  • 反复正反向扩散在可控噪声下探测模型深层记忆区域。
  • 深度吸引子可经数千次扰动后复原,包含训练集中的真实图片。
  • 仅需采样控制,适用于隐私审计与版权检测,通用性强。

我们提出循环去噪——在可控噪声水平下重复正向与逆向扩散过程——作为图像扩散模型的提取攻击方法。受无序固体中随机组织的启发,该方法揭示了标准采样难以触及的模型学习分布区域。动态过程引导样本趋向具有宽稳定性谱的吸引子,其中最深的吸引子为超稳定态:即使遭受近乎完全破坏,仍可在数千次去噪-加噪循环后复现。这些吸引子多数对应训练数据中的真实图像,包括商业照片、品牌水印及网络爬取痕迹。该攻击仅需采样层控制,无需梯度、权重查看、提示词、字幕或训练数据先验知识。与依赖大规模生成和事后过滤的生成-筛选攻击不同,本方法为完全无条件协议。我们在 Stable Diffusion v1.4 和像素空间 DDPM 中均验证该现象,显示潜空间与像素空间模型行为一致。在不同噪声幅度下观察到类似屈服的相变:低噪声引发平凡吸收固定点或极限环,高噪声则诱发重组、势阱跳跃及在结构化记忆吸引子盆地中的长时捕获。此外还发现层级部分吸收、提示词稳定势阱以及跨初始条件的吸引子集合普适性。结果表明,循环去噪既是基于物理的生成景观探测工具,也是实用的记忆审计手段,对隐私保护、版权合规与模型指纹识别具有重要意义。

原文摘要 · Abstract (English)

We introduce cyclic denoising -- repeated forward and reverse diffusion at controlled noise amplitudes -- as an extraction attack for image diffusion models. Inspired by random organization in disordered solids, cyclic denoising exposes regions of the learned distribution that are largely inaccessible to standard sampling. The dynamics drive samples toward attractors with a broad stability spectrum. The deepest attractors are ultrastable: they regenerate after near-total corruption and persist through thousands of noising-denoising cycles. Many of these attractors correspond to memorized training images, including stock photographs, brand watermarks, and web-crawl artifacts. The attack requires only sampler-level control, with no gradients, weight inspection, prompts, captions, or prior knowledge of the training data. Unlike generate-and-filter attacks, which rely on large-scale prompted generation and post-hoc similarity or membership-inference filtering, our main protocol is fully unconditioned. We demonstrate the phenomenon in Stable Diffusion v1.4 and in a pixel-space DDPM, showing consistent behavior across latent- and pixel-space diffusion models. Across noise amplitudes, we observe a yielding-like transition: low-amplitude cycling produces trivial absorbing fixed points or limit cycles, while larger amplitudes induce rearrangements, basin hopping, and long-lived trapping in structured memorized attractor basins. We also observe hierarchical partial absorption, prompt-stabilized basins, and cross-initial-condition universality of the recovered attractor set. Our results therefore show that cyclic denoising is both a physics-inspired probe of generative landscapes and a practical tool for memorization auditing, with implications for privacy, copyright compliance, and model fingerprinting.

扩散模型记忆提取隐私安全模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。