arXiv:2510.08091cs.CLcs.AI2025-10ACL被引 1

LLM生成的推理能显著影响人类和模型对常识答案的可信度判断。

Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility

  • 用LLM生成正反论证,测试其对判断的影响。
  • 人类对答案可信度评分随论证方向显著升高或降低。
  • 揭示了LLM对认知的潜在操控力,适合研究者与AI伦理关注者阅读。

我们研究了人类(及大语言模型)在多项选择常识基准题中,对答案可信度的判断在面对(不)合理论证时受何种影响,尤其聚焦于由大语言模型生成的推理过程。我们收集了3000条人类判断和13600条大语言模型的判断。结果显示,在存在大语言模型生成的正向论证(PRO)时,人类平均可信度评分上升;在存在反向论证(CON)时则下降,表明人类普遍认为这些论证具有说服力。对大语言模型的实验也显示出类似影响模式。研究结果展示了大语言模型在探究人类认知方面的创新用途,同时也引发担忧:即使在人类被视为‘专家’的常识领域,大语言模型仍可能显著影响人们的信念。

原文摘要 · Abstract (English)

We investigate the degree to which human (and LLM) plausibility judgments of multiple-choice commonsense benchmark answers are subject to influence by (im)plausibility arguments for or against an answer, in particular, using rationales generated by LLMs. We collect 3,000 plausibility judgments from humans and another 13,600 judgments from LLMs. Overall, we observe increases and decreases in mean human plausibility ratings in the presence of LLM-generated PRO and CON rationales, respectively, suggesting that, on the whole, human judges find these rationales convincing. Experiments with LLMs reveal similar patterns of influence. Our findings demonstrate a novel use of LLMs for studying aspects of human cognition, while also raising practical concerns that, even in domains where humans are ``experts'' (i.e., common sense), LLMs have the potential to exert considerable influence on people's beliefs.

认知影响常识推理LLM可信度人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。