arXiv:2509.00488cs.LG2025-09中稿 · ICML被引 2

定位图像生成模型的隐私泄露点,针对性降低记忆风险。

Localizing and Mitigating Memorization in Image Autoregressive Models

  • 按分辨率层级分析记忆分布,发现高层级更易产生记忆
  • 干预高记忆组件可大幅减少数据泄露,不影响生成质量
  • 适合关注生成模型隐私安全的研究者和开发者

图像自回归(IAR)模型在生成速度与质量上已达顶尖水平,但其训练数据的记忆问题引发隐私担忧。本文通过细粒度测量,探究不同IAR架构中记忆现象的发生位置与方式。结果表明:在分层分辨率架构中,记忆早期出现并随分辨率加深;而在标准逐标记自回归模型中,记忆集中于后期处理阶段。这些记忆模式与数据泄露能力密切相关。通过对高记忆组件进行干预,可显著降低数据提取能力,且对生成图像质量影响极小。研究揭示了图像生成模型内部行为机制,为缓解隐私风险提供了可行策略。

原文摘要 · Abstract (English)

Image AutoRegressive (IAR) models have achieved state-of-the-art performance in speed and quality of generated images. However, they also raise concerns about memorization of their training data and its implications for privacy. This work explores where and how such memorization occurs within different image autoregressive architectures by measuring a fine-grained memorization. The analysis reveals that memorization patterns differ across various architectures of IARs. In hierarchical per-resolution architectures, it tends to emerge early and deepen with resolutions, while in IARs with standard autoregressive per token prediction, it concentrates in later processing stages. These localization of memorization patterns are further connected to IARs' ability to memorize and leak training data. By intervening on their most memorizing components, we significantly reduce the capacity for data extraction from IARs with minimal impact on the quality of generated images. These findings offer new insights into the internal behavior of image generative models and point toward practical strategies for mitigating privacy risks.

图像生成隐私安全自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。