arXiv:2601.13433cs.CLcs.LG2026-01被引 5

模型更易受权威误导,且越权威越自信犯错。

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models

  • 用四种专业角色测试模型对权威建议的响应
  • 专家级误导使模型准确率下降,错误答案更自信
  • 偏差可被修正,即使专家给出错误建议

先前研究显示语言模型在推理任务中的表现会受到提示和建议的影响。然而,推荐来源可信度的影响仍缺乏深入探讨。本文在数学、法律和医学三个领域的4个数据集上,评估了11个模型对四种不同专业水平角色提供的推荐的响应。结果表明,随着推荐者权威性提升,模型对错误或误导性建议的敏感度显著上升,不仅导致准确率下降,还使其对错误答案产生更高的置信度。我们进一步证明这种权威偏差是模型内部机制所编码的,可通过干预手段消除偏差,从而在专家提供错误建议时仍能提升模型表现。

原文摘要 · Abstract (English)

Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However, the influence of endorsement source credibility remains underexplored. We investigate whether language models exhibit systematic bias based on the perceived expertise of the provider of the endorsement. Across 4 datasets spanning mathematical, legal, and medical reasoning, we evaluate 11 models using personas representing four expertise levels per domain. Our results reveal that models are increasingly susceptible to incorrect/misleading endorsements as source expertise increases, with higher-authority sources inducing not only accuracy degradation but also increased confidence in wrong answers. We also show that this authority bias is mechanistically encoded within the model and a model can be steered away from the bias, thereby improving its performance even when an expert gives a misleading endorsement.

权威偏差语言模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。