大模型在评判语言得体性上表现好,但自己说时却常不得体。
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
- 对比模型作为语言听众和说话人的表现差异
- 多数模型听懂了却说不好,评判能力远超生成能力
- 提醒研究者需更综合评估模型的语用能力
大型语言模型(LLMs)被广泛视为语言知识的存储库。当前研究通常将模型同时用作语言生成者和语言评判者,但这两个角色很少被直接关联考察。因此,二者表现是否一致尚不明确。本文通过比较多个开源与专有模型在三种语用场景下的表现,分别评估其作为语用听众(判断语言输出是否恰当)和语用说话人(生成恰当语言)的能力。结果显示,模型在语用评判上的表现显著优于语用生成,存在稳健的不对称性。这表明当前大模型在语用判断与生成之间仅弱度对齐,呼吁采用更整合的评估方法。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly studied as repositories of linguistic knowledge. In this line of work, models are commonly evaluated both as generators of language and as judges of linguistic output, yet these two roles are rarely examined in direct relation to one another. As a result, it remains unclear whether success in one role aligns with success in the other. In this paper, we address this question for pragmatic competence by comparing LLMs' performance as pragmatic listeners, judging the appropriateness of linguistic outputs, and as pragmatic speakers, generating pragmatically appropriate language. We evaluate multiple open-weight and proprietary LLMs across three pragmatic settings. We find a robust asymmetry between pragmatic evaluation and pragmatic generation: many models perform substantially better as listeners than as speakers. Our results suggest that pragmatic judging and pragmatic generation are only weakly aligned in current LLMs, calling for more integrated evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。