让大模型生成更公平,同一问题不同用户得到同样高质量回答。
Identity-Robust Language Model Generation via Content Integrity Preservation
- 不训练、不修改模型,仅通过提示词过滤无关身份信息。
- 在18种社会身份下,偏差平均降低77%。
- 适合关注生成结果公平性的应用开发者和评测人员。
大语言模型的输出常因用户社会人口属性差异而变化,导致事实准确性、实用性和安全性下降,即使在与身份无关的客观问题上也存在此类现象。不同于以往对刻板印象或表征偏差的研究,本文聚焦于核心响应质量的身份依赖性退化。实证发现,尽管事实知识在各类身份中均被稳健编码,但生成行为存在偏见。针对这一不匹配,我们提出一种轻量级、无需训练的鲁棒生成框架,通过选择性中和非关键身份信息,同时保留语义必要属性,从而维持内容完整性。在四个基准测试和18种社会人口身份上的实验表明,相比原始提示,身份相关偏差平均降低77%;相较于提示增强防御方法,降低45%。本工作填补了缓解提示中用户身份线索对核心生成质量影响的关键空白。
原文摘要 · Abstract (English)
Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior work on stereotypical or representational bias, this paper studies identity-dependent degradation of core response quality. We show empirically that such degradation arises from biased generation behavior, despite factual knowledge being robustly encoded across identities. Motivated by this mismatch, we propose a lightweight, training-free framework for identity-robust generation that selectively neutralizes non-critical identity information while preserving semantically essential attributes, thus maintaining output content integrity. Experiments across four benchmarks and 18 sociodemographic identities demonstrate an average 77% reduction in identity-dependent bias compared to vanilla prompting and a 45% reduction relative to prompt-based defenses. Our work addresses a critical gap in mitigating the impact of user identity cues in prompts on core generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。