arXiv:2503.20098cs.LG2025-03中稿 · AISTATS 2025被引 5

从信息论角度揭示概念擦除的极限,实现无损擦除。

Fundamental Limits of Perfect Concept Erasure

  • 用信息论建模概念擦除的理论边界
  • 实证验证擦除函数达理论最优
  • 适合需要高精度公平性的模型开发者

概念擦除旨在从表示集中移除特定概念(如性别或种族)的信息,同时尽可能保留原始表示的有用性。该任务在实现公平性和理解模型性能影响方面具有重要应用。以往方法更关注概念擦除的鲁棒性,但擦除与保留效用之间存在固有权衡,难以实现完美擦除且保持高实用性。本文从信息论视角出发,量化概念擦除的根本极限,分析实现完美擦除所需的数据分布和擦除函数的约束条件。实验表明,所推导的擦除函数可达到理论最优边界,并在一系列合成与真实数据集(使用GPT-4表示)上优于现有方法。

原文摘要 · Abstract (English)

Concept erasure is the task of erasing information about a concept (e.g., gender or race) from a representation set while retaining the maximum possible utility -- information from original representations. Concept erasure is useful in several applications, such as removing sensitive concepts to achieve fairness and interpreting the impact of specific concepts on a model's performance. Previous concept erasure techniques have prioritized robustly erasing concepts over retaining the utility of the resultant representations. However, there seems to be an inherent tradeoff between erasure and retaining utility, making it unclear how to achieve perfect concept erasure while maintaining high utility. In this paper, we offer a fresh perspective toward solving this problem by quantifying the fundamental limits of concept erasure through an information-theoretic lens. Using these results, we investigate constraints on the data distribution and the erasure functions required to achieve the limits of perfect concept erasure. Empirically, we show that the derived erasure functions achieve the optimal theoretical bounds. Additionally, we show that our approach outperforms existing methods on a range of synthetic and real-world datasets using GPT-4 representations.

概念擦除信息论公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。