arXiv:2510.23650cs.LGcs.AI2025-10
通过语义感知的输出层干预,有效降低大模型偏见且保持语言流畅性。
Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs
- 在输出层直接干预,无需修改隐藏层参数。
- 动态方法可减少70%偏见,同时几乎不损失语言质量。
- 适合希望快速提升对齐模型公平性的研究者与开发者。
我们提出了两种零样本的输出层去偏方法:静态(Static)与动态(Dynamic)。动态方法在保持语言流畅性几乎不变的前提下,将偏见降低高达70%。实验表明,输出层干预优于传统隐藏层修改方法。语义感知的输出层干预在对齐大模型上表现出稳定且高效的去偏效果。
原文摘要 · Abstract (English)
We proposed Static and Dynamic -- two zero-shot logits-layer debiasing methods. Dynamic reduces bias by up to 70% with minimal fluency loss. Logits intervention outperforms hidden-layer approaches. We show semantic-aware logits intervention is stable and effective for debiasing aligned LLMs.
去偏大模型输出层干预
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。