arXiv:2601.10896cs.CL2026-01ACL被引 1

发现大模型对话评价会因表述方式不同而偏袒说话人,提出检测与修正框架。

DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

  • 通过对比陈述句与归因句的判断差异,识别出模型对发言者的无意识偏袒
  • 在10个领域3000+样本中,平均判断偏差达15.9个百分点,但准确率几乎不变
  • 适合评估大模型公平性、用于人机交互系统设计或可信评测体系构建

大模型被广泛用作第三方评判者,但其在对话场景下的评价可靠性仍不清楚。我们发现,相同内容因表述方式不同,获得的评价结果显著不同:当内容以待验证陈述形式呈现时,模型判断更客观;而当内容被归因于某说话人时,模型更易产生偏袒。这种现象称为对话性盲从(Dialogic Deference)。本文提出DialDefer框架,引入对话性盲从分数(DDS)来量化此类方向性偏差。在十个领域、超过3000个样本及五种模型上,对话框架引发的平均绝对偏差为15.9个百分点(p < .0001),而准确率变化不足2个百分点。在自然语言的Reddit对话中,该效应放大2至5倍。该现象具有领域依赖性:同一模型在研究生级科学议题上更倾向质疑,在社会判断任务上则更倾向服从。消融实验表明,人类与大模型的归属差异是主要驱动因素(偏差达17.7个百分点),暗示模型将挑战人类观点视为更高成本。尽管可尝试缓解盲从,但过度纠正会导致转向过度怀疑,暴露出校准问题远超单纯提升准确率。

原文摘要 · Abstract (English)

LLMs are increasingly used as third-party judges, yet their reliability when evaluating speakers in dialogue remains poorly understood. We show that LLMs judge identical claims differently depending on framing: the same content receives different verdicts when presented as a statement to verify ("Is this statement correct?") versus attributed to a speaker ("Is this speaker correct?"). We call this dialogic deference and introduce DialDefer, a framework for detecting and mitigating these framing-induced judgment shifts. Our Dialogic Deference Score (DDS) captures directional shifts that aggregate accuracy obscures. Across ten domains, 3k+ instances, and five models, conversational framing induces large shifts (mean|DDS|=15.9 percentage points (pp) across models, p < .0001) while accuracy remains stable (<2 pp), with effects amplifying 2--5x on naturalistic Reddit conversations. This effect is domain-dependent: a single model can shift toward disagreement (skepticism) on graduate-level science and toward agreement (deference) on social judgment. Ablations reveal that human-vs-LLM attribution drives the largest shifts (17.7 pp swing), suggesting models treat disagreement with humans as more costly than with AI. Mitigation attempts can reduce deference but over-correct into skepticism, revealing a calibration problem beyond accuracy optimization.

大模型评测认知偏见对话系统信任机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。