模型输出公平但内部有偏,重置隐藏信息可逆转决策。
Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions

- 通过激活控制和跨层干预,发现隐藏偏见影响决策
- 重置关键层信息后,决策几乎完全反转
- 偏见效应不对称,且易被提示或微调操控
指令微调的语言模型在高风险决策中表现出行为公平性,但在内部表示中仍保留偏见关联。我们使用带有公开权重的模型进行房贷审批研究,采用仅在姓名上具有种族关联差异的匹配申请数据集。结果显示:模型输出无偏差,但其各层表示中仍存在并放大了人口特征信息。通过激活控制和新颖的跨层干预手段,我们证明这些被抑制的信息具有决策相关性:在关键层重新注入后,可导致近乎完全的决策反转。关键的是,这种潜在偏见具有不对称性——控制干预只在某一方向显著影响决策,反向影响极小,且对对抗性提示工程和参数高效微调敏感。这表明,仅关注输出的公平性审计是不够的,公平输出可能掩盖可被利用的内部偏见。研究呼吁建立结合输出评估与表征分析的双层测试框架,以提升高风险决策中AI治理的有效性。
原文摘要 · Abstract (English)
Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representations. However, whether these suppressed representations can affect model outputs - and whether such causal potency is symmetric across demographic groups - remains unknown. We investigate the use of open-weight models for mortgage underwriting using matched applications that differ only in racially-associated names and reveal a critical disconnect: models show no output-level bias, yet retain and amplify demographic representations across model layers. Through activation steering and novel cross-layer interventions, we demonstrate that this suppressed information is decision-relevant: when reinjected at critical layers, it produces near-complete decision reversals. Critically, this latent bias is asymmetric - steering interventions affect decisions in one demographic direction, while producing minimal effects in reverse - and susceptible to adversarial prompt engineering and parameter-efficient fine-tuning. These findings demonstrate that behavioural audits focused on outputs are insufficient: fair outputs can mask exploitable internal biases. They also motivate dual-layer testing frameworks combining output evaluation with representational analysis for AI governance in high-stakes decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。