arXiv:2510.19687cs.CLcs.AI2025-10NeurIPS被引 8

研究大模型能否识别说话动机,发现其具备基本敏感性但需改进。

Are Large Language Models Sensitive to the Motives Behind Communication?

  • 用认知科学实验验证大模型能像人一样评估有动机的信息。
  • 在广告场景中,模型对动机的判断与理性模型偏差较大。
  • 强化动机提示可显著提升模型判断准确性,适合做可信度评估研究者参考。

人类交流具有意图:人们在表达时往往带有特定动机。因此,大型语言模型(LLMs)处理的信息天然受制于人的意图与激励。人类擅长识别这些细微差别,如辨别善意或利己动机以决定信任与否。为使大模型在真实世界中有效运作,它们也需评估信息源的动机,例如权衡销售宣传的可信度。本文系统研究了大模型是否具备这种动机警觉能力。首先,通过认知科学中的受控实验,验证大模型的行为符合理性学习模型,能以类人方式降低对偏见来源信息的信任。接着,将评估扩展至赞助式在线广告这一更自然的场景,发现大模型的推断与理性模型预测的契合度明显下降,部分因额外干扰信息分散了对动机的关注。然而,通过简单干预增强动机和激励的显著性,显著提升了大模型与理性模型的一致性。结果表明,大模型具备基础动机敏感性,但在复杂现实场景中推广仍需进一步优化。

原文摘要 · Abstract (English)

Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs) and AI agents process is inherently framed by humans' intentions and incentives. People are adept at navigating such nuanced information: we routinely identify benevolent or self-serving motives in order to decide what statements to trust. For LLMs to be effective in the real world, they too must critically evaluate content by factoring in the motivations of the source -- for instance, weighing the credibility of claims made in a sales pitch. In this paper, we undertake a comprehensive study of whether LLMs have this capacity for motivational vigilance. We first employ controlled experiments from cognitive science to verify that LLMs' behavior is consistent with rational models of learning from motivated testimony, and find they successfully discount information from biased sources in a human-like manner. We then extend our evaluation to sponsored online adverts, a more naturalistic reflection of LLM agents' information ecosystems. In these settings, we find that LLMs' inferences do not track the rational models' predictions nearly as closely -- partly due to additional information that distracts them from vigilance-relevant considerations. However, a simple steering intervention that boosts the salience of intentions and incentives substantially increases the correspondence between LLMs and the rational model. These results suggest that LLMs possess a basic sensitivity to the motivations of others, but generalizing to novel real-world settings will require further improvements to these models.

动机识别可信度评估大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。