arXiv:2602.00199cs.LG2026-02

通过黎曼贝叶斯推断减少生成模型的过拟合记忆。

Reducing Memorisation in Generative Models via Riemannian Bayesian Inference

  • 用黎曼度量捕捉损失函数几何结构,构建自适应后验分布。
  • 在保持生成能力的同时显著降低模型对训练数据的记忆。
  • 适合关注生成模型泛化与隐私保护的研究者。

现代生成模型可生成逼真样本,但如何平衡记忆与泛化仍是开放问题。本文从贝叶斯视角出发,聚焦流匹配与扩散模型的参数空间,构建能更好捕捉数据分布变异性、更优的预测后验。特别地,利用黎曼度量刻画损失函数几何结构,并采用灵活的近似后验以适配损失景观的局部结构。该方法可采样出与原模型相似但记忆性更低的生成模型。实验表明,所提方法在保持泛化能力的同时有效降低记忆。此外,我们提供了理论分析,解释了实验结果。整体而言,本工作表明考虑损失几何结构可有效利用复杂高维生成模型的参数空间。

原文摘要 · Abstract (English)

Modern generative models can produce realistic samples, however, balancing memorisation and generalisation remains an open problem. We approach this challenge from a Bayesian perspective by focusing on the parameter space of flow matching and diffusion models and constructing a predictive posterior that better captures the variability of the data distribution. In particular, we capture the geometry of the loss using a Riemannian metric and leverage a flexible approximate posterior that adapts to the local structure of the loss landscape. This approach allows us to sample generative models that resemble the original model, but exhibit reduced memorisation. Empirically, we demonstrate that the proposed approach reduces memorisation while preserving generalisation. Further, we provide a theoretical analysis of our method, which explains our findings. Overall, our work illustrates how considering the geometry of the loss enables effective use of the parameter space, even for complex high-dimensional generative models.

生成模型贝叶斯推断过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。