arXiv:2603.18530cs.CLcs.AI2026-03被引 1

发现大模型决策受身份、权威和表述方式影响,提出检测与缓解系统性偏见的方法。

When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making

  • 通过干预一致性测试,识别模型对身份、权威和表述的虚假依赖。
  • 权威偏差(5.8%)和表述偏差(5.0%)远高于种族偏差(2.2%),金融领域尤为严重。
  • 设计可迭代优化的检测-诊断-缓解闭环,实现78%偏见降低,适合高风险应用审查。

大型语言模型(LLMs)被广泛用于高风险决策,但其对虚假特征的敏感性尚未充分揭示。我们提出ICE-Guard框架,通过干预一致性测试检测三类虚假特征依赖:人口属性(姓名/种族替换)、权威性(资质/声誉替换)和表述方式(正负表述重述)。在涵盖10个高风险领域的3000个情景中,评估了来自8个模型家族的11个LLM,发现:(1)权威偏差(均值5.8%)和表述偏差(5.0%)显著高于人口偏差(2.2%),挑战了学界对人口因素的过度聚焦;(2)偏差集中在特定领域——金融领域权威偏差达22.6%,而刑事司法仅2.8%;(3)结构化分解方法(模型提取特征+确定性规则判断)使翻转率降低最多100%(9个模型中位数降幅49%)。我们演示了基于ICE的检测-诊断-缓解-验证循环,通过迭代提示修补实现累计78%偏见减少。与真实COMPAS再犯数据对比显示,基于合成数据的翻转率低于真实数据,表明本基准提供了保守估计。代码与数据已公开。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for high-stakes decisions, yet their susceptibility to spurious features remains poorly characterized. We introduce ICE-Guard, a framework applying intervention consistency testing to detect three types of spurious feature reliance: demographic (name/race swaps), authority (credential/prestige swaps), and framing (positive/negative restatements). Across 3,000 vignettes spanning 10 high-stakes domains, we evaluate 11 LLMs from 8 families and find that (1) authority bias (mean 5.8%) and framing bias (5.0%) substantially exceed demographic bias (2.2%), challenging the field's narrow focus on demographics; (2) bias concentrates in specific domains -- finance shows 22.6% authority bias while criminal justice shows only 2.8%; (3) structured decomposition, where the LLM extracts features and a deterministic rubric decides, reduces flip rates by up to 100% (median 49% across 9 models). We demonstrate an ICE-guided detect-diagnose-mitigate-verify loop achieving cumulative 78% bias reduction via iterative prompt patching. Validation against real COMPAS recidivism data shows COMPAS-derived flip rates exceed pooled synthetic rates, suggesting our benchmark provides a conservative estimate of real-world bias. Code and data are publicly available.

大模型偏见决策公平性干预测试可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。