arXiv:2509.06952cs.CL2025-09EMNLP被引 6

用沟通游戏评测大模型的语用推理能力,发现大模型在理解上接近人类,生成上可借助思维链和贝叶斯推理提升。

On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts

  • 基于Wavelength游戏框架,评估模型在理解和生成中的语用推理能力。
  • 大模型在理解任务中达到接近人类的准确率,无需思维链或贝叶斯推理即可与人类判断高度相关。
  • 使用贝叶斯语用推理(RSA)显著提升生成表现,尤其在复杂概念沟通中优势明显。

语言使用受语用学影响——即在具体情境中对交际目标和规范的推理。随着语言模型被广泛用于对话代理,理解其语用推理能力变得愈发重要。本文提出一个基于Wavelength通信游戏的评估框架,该游戏要求说话者与听者就广泛概念进行细致沟通。我们在多种语言模型上测试了语言理解与生成能力,采用直接提示和思维链(CoT) prompting,并进一步探索将贝叶斯语用推理(RSA)融入模型推理的方法。结果表明,先进语言模型在理解任务中表现优异,准确率接近人类,且与人类判断高度相关,甚至无需思维链或RSA。在生成任务中,思维链优于直接提示,而使用RSA则显著优于两者。研究揭示了语言模型在语用推理中的强项与局限,展示了通过RSA提升其能力的潜力,为理解模型与人类的概念表征、语言理解及社会推理提供了新方向。

原文摘要 · Abstract (English)

Language use is shaped by pragmatics -- i.e., reasoning about communicative goals and norms in context. As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities. We propose an evaluation framework derived from Wavelength, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner. We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference. We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA. On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches. Our study helps identify the strengths and limitations in LMs' pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans.

语用推理语言模型贝叶斯推理对话评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。