用自述文本隐式消除模型偏见,无需敏感属性信息
Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles
- 通过用户自述文本提取隐式偏见信号,实现后处理去偏
- 在多个语言模型上,隐式方法效果不逊于甚至优于显式标签方法
- 提出新基准,适合研究真实场景下无敏感属性的去偏问题
大多数NLP公平性研究依赖性别、种族等敏感属性的直接访问,但现实中这些信息常因隐私、缺失元数据或法律限制而不可得,尽管模型仍能从文本中推断。这引出核心问题:能否在无敏感属性的情况下实现有效去偏?本文提出H-SAL方法,利用自述文本作为隐式信号,进行后处理的概念与属性擦除。为此,我们构建了一个基于多领域Stack Exchange的公平性基准,用于帮助性预测任务,包含显式与隐式信号,支持对比有无敏感标签的去偏效果。在编码器与解码器型语言模型上,实验表明隐式自述文本的表现往往与显式标签相当甚至更优。结果拓展了表示层面公平性研究,并为真实数据约束下的去偏提供了新基准。
原文摘要 · Abstract (English)
Most fairness research in NLP assumes direct access to protected attributes such as gender, race, or nationality. In practice, however, such information is often unavailable due to privacy constraints, missing metadata, or legal restrictions, even though models may infer it from indirect textual cues. This raises a key question: can debiasing succeed without direct access to sensitive attributes? We propose H-SAL, which performs post-hoc concept and attribute erasure using self-description text as an implicit debiasing signal. To support this setting, we introduce a multi-domain Stack Exchange-based fairness benchmark for helpfulness prediction that includes both explicit and implicit signals, enabling comparison between standard debiasing with protected labels and debiasing without access to sensitive information. Across encoder and decoder-only language models, we find that implicit self-description often matches or outperforms explicit-label-based debiasing. Our results broaden representation-level fairness research and provide a new benchmark for studying debiasing under realistic data constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。