通过近邻密度估计实现数据去相关,提升公平性与隐私保护效果
Nearest-Neighbor Density Estimation for Dependency Suppression
- 基于变分自编码器与近邻密度估计设计新损失函数
- 在多个数据集上优于现有无监督方法,媲美有监督方法
- 适合需要消除敏感属性影响的公平学习与隐私保护场景
消除数据中不想要的依赖关系在公平性、鲁棒学习和隐私保护等领域至关重要。本文提出一种基于编码器的方法,学习与敏感变量无关但保留数据核心特征的表示。不同于依赖去相关或对抗学习的现有方法,本方法显式估计并调整数据分布以消除统计依赖。通过结合专用变分自编码器与基于非参数近邻密度估计的新损失函数,实现独立性的直接优化。在多个数据集上的实验表明,该方法可超越现有无监督技术,甚至在信息去除与数据效用平衡方面媲美有监督方法。
原文摘要 · Abstract (English)
The ability to remove unwanted dependencies from data is crucial in various domains, including fairness, robust learning, and privacy protection. In this work, we propose an encoder-based approach that learns a representation independent of a sensitive variable but otherwise preserving essential data characteristics. Unlike existing methods that rely on decorrelation or adversarial learning, our approach explicitly estimates and modifies the data distribution to neutralize statistical dependencies. To achieve this, we combine a specialized variational autoencoder with a novel loss function driven by non-parametric nearest-neighbor density estimation, enabling direct optimization of independence. We evaluate our approach on multiple datasets, demonstrating that it can outperform existing unsupervised techniques and even rival supervised methods in balancing information removal and utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。