arXiv:2411.01562cs.CLcs.AI2024-11被引 7

测试大模型是否像人一样会根据语境合理说话

Are LLMs good pragmatic speakers?

  • 用人类语用推理框架对比大模型与人类的表达选择
  • 大模型与人类在语用评分上相关性较弱,未达共情水平
  • 为后续改进大模型语用能力提供研究方向

大型语言模型(LLMs)在被认为包含自然语言语用信息的数据上训练,但它们是否真正表现出语用说话者的行为?本文借助理性言语行为(Rational Speech Act, RSA)框架,该框架模拟人类交流中的语用推理。基于TUNA语料库构建的指称游戏范式,我们对先进大模型Llama3-8B-Instruct与RSA模型生成的候选指称语句进行评分,并比较两者表现。由于RSA需要定义替代表达和真值条件语义函数,我们探索了不同设定下的对比结果。发现尽管大模型评分与RSA存在部分正相关,但尚无足够证据表明其行为符合语用说话者特征。本研究为未来针对不同模型与设置的深入探索(包括人类受试者评估)提供了基础,以判断大模型能否或如何真正具备语用说话者行为。

原文摘要 · Abstract (English)

Large language models (LLMs) are trained on data assumed to include natural language pragmatics, but do they actually behave like pragmatic speakers? We attempt to answer this question using the Rational Speech Act (RSA) framework, which models pragmatic reasoning in human communication. Using the paradigm of a reference game constructed from the TUNA corpus, we score candidate referential utterances in both a state-of-the-art LLM (Llama3-8B-Instruct) and in the RSA model, comparing and contrasting these scores. Given that RSA requires defining alternative utterances and a truth-conditional meaning function, we explore such comparison for different choices of each of these requirements. We find that while scores from the LLM have some positive correlation with those from RSA, there isn't sufficient evidence to claim that it behaves like a pragmatic speaker. This initial study paves way for further targeted efforts exploring different models and settings, including human-subject evaluation, to see if LLMs truly can, or be made to, behave like pragmatic speakers.

大模型语用学自然语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。