arXiv:2608.22887cs.AIcs.CL2026-08

发现大模型决策中代理依赖与真实证据不匹配,易误判为歧视

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

  • 通过临床排序任务量化模型对代理变量的真实依赖程度
  • 三类结论:过度依赖、合理依赖、依赖不足,且普遍存在低估证据现象
  • 社会标签虽抑制依赖但极易被上下文例子突破,适合模型审计研究者

大型语言模型正被用于分诊和贷款等关键决策场景,需区分合理推断与非法代理使用。现有审计方法仅关注改变人口属性时决策是否变化,但相关属性本身具有预测价值,因此决策变化未必是歧视。本文在具备已知真实答案的临床排序任务中,测量了四种LLM的因果代理效应,精确计算出依据证据应有之依赖度作为参考基准。单一审计信号可得出三种结论:过度依赖、合理依赖、依赖不足。在中性标签下,所有模型均对无信息的代理变量产生依赖;当代理变量具信息时,三类情况均出现。社会领域名称会降低依赖程度,但在一个模型中低于参考值。两个原因解释此现象:依赖严重滞后于证据,且社会标签抑制机制脆弱——上下文示例使所有模型的依赖度回升至零以上。基于准确率的评估完全无法捕捉这些现象。

原文摘要 · Abstract (English)

Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.

大模型审计代理依赖公平性临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。