在移除敏感概念时,精准保留与目标任务的关联性。
Preserving Task-Relevant Information Under Linear Concept Removal
- 通过斜投影方法精准剥离无关概念方向。
- 在移除敏感属性的同时,完全保留其与目标标签的协方差。
- 适合需要公平性与信息保真的场景,如医疗、招聘模型
现代神经网络常将无关概念与任务相关信息一同编码,引发公平性和可解释性问题。现有后处理方法虽能消除不希望的概念,但常损害有用信号。本文提出SPLINCE——一种用于线性概念移除与协方差保持的同步投影方法,可在消除敏感概念的同时,精确保留其与目标标签的协方差。SPLINCE通过斜投影实现‘剪切’无用方向,同时保护关键标签相关性。理论上,它是唯一能消除线性概念可预测性并最小化嵌入失真的解。实验上,SPLINCE在Bias in Bios和Winobias等基准上优于基线,有效去除受保护属性,且对主任务信息损伤极小。
原文摘要 · Abstract (English)
Modern neural networks often encode unwanted concepts alongside task-relevant information, leading to fairness and interpretability concerns. Existing post-hoc approaches can remove undesired concepts but often degrade useful signals. We introduce SPLINCE-Simultaneous Projection for LINear concept removal and Covariance prEservation - which eliminates sensitive concepts from representations while exactly preserving their covariance with a target label. SPLINCE achieves this via an oblique projection that 'splices out' the unwanted direction yet protects important label correlations. Theoretically, it is the unique solution that removes linear concept predictability and maintains target covariance with minimal embedding distortion. Empirically, SPLINCE outperforms baselines on benchmarks such as Bias in Bios and Winobias, removing protected attributes while minimally damaging main-task information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。