arXiv:2602.13568cs.AI2026-02被引 3

大模型更信任人类专家而非其他大模型的判断。

Who Do LLMs Trust? Human Experts Matter More Than Other LLMs

  • 通过设定不同来源的建议,测试大模型的决策倾向。
  • 模型在多数任务中更倾向采纳标注为专家的人类意见,即使错误也调整答案。
  • 适合关注大模型社会性推理与可信源影响的研究者阅读。

大型语言模型(LLMs)在运行环境中常接触其他智能体的回答、工具输出或人类建议。在人类中,此类信息的影响取决于来源可信度和共识强度。本文研究了大模型是否表现出类似的社会影响模式,并探究其是否更倾向于接受人类反馈而非其他大模型的反馈。在阅读理解、多步推理和道德判断三个二元决策任务中,我们向四款指令微调的LLM提供前序回答,这些回答被标记为来自朋友、人类专家或其它大模型。我们操控群体正确性并改变群体规模。第二项实验中引入单个人类与单个大模型之间的直接分歧。结果表明,模型显著更倾向于采纳标注为人类专家的意见,包括该信号错误时;且对专家意见的修正速度明显快于对其他大模型意见的修正。这些发现揭示,专家标签对当前大模型构成强先验,体现出跨决策领域泛化的可信度敏感社会影响。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly operate in environments where they encounter social information such as other agents' answers, tool outputs, or human recommendations. In humans, such inputs influence judgments in ways that depend on the source's credibility and the strength of consensus. This paper investigates whether LLMs exhibit analogous patterns of influence and whether they privilege feedback from humans over feedback from other LLMs. Across three binary decision-making tasks, reading comprehension, multi-step reasoning, and moral judgment, we present four instruction-tuned LLMs with prior responses attributed either to friends, to human experts, or to other LLMs. We manipulate whether the group is correct and vary the group size. In a second experiment, we introduce direct disagreement between a single human and a single LLM. Across tasks, models conform significantly more to responses labeled as coming from human experts, including when that signal is incorrect, and revise their answers toward experts more readily than toward other LLMs. These results reveal that expert framing acts as a strong prior for contemporary LLMs, suggesting a form of credibility-sensitive social influence that generalizes across decision domains.

大模型社会影响可信源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。