用标量多样性测试大模型语用推理,发现评估方式影响结果
Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity

- 通过标量多样性测试不同模型的语用推理能力
- 两种评估方法表现不一,无明显优劣
- 模型家族与提示策略共同影响推理效果
评估大语言模型(LLM)的语用推理能力仍具挑战性,因模型行为受评估方法影响。以往研究指出,基于提示的判断可能偏离模型内部概率分布,引发对性能是否反映真实能力还是任务诱导行为的疑问。本研究采用标量多样性作为分级诊断工具,比较直接概率测量与元语言提示在多个模型和实验设置下的表现。结果显示,两种方法无一致优势,且语用行为在不同模型族、提示策略与任务结构间差异显著。仅特定模型-条件组合中出现标量多样性梯度,表明大模型的语用推理是内部概率表征与任务诱导提示行为的交互结果,而非单一评估范式能稳定捕捉的固有能力。该发现强调评估设计在解读大模型语用能力中的核心作用。
原文摘要 · Abstract (English)
Evaluating pragmatic reasoning in large language models (LLMs) remains challenging because model behavior can vary depending on evaluation methods. Previous studies suggest that prompt-based judgments may diverge from models' internal probability distributions, raising questions about whether observed performance reflects underlying competence or task-induced behavior. This study examines this issue using scalar diversity as a graded diagnostic for pragmatic inference. Following Hu & Levy (2023), this study compares direct probability measurement and metalinguistic prompting across multiple models and experimental settings. The results show that neither evaluation method consistently outperforms the other and that pragmatic behavior varies substantially across model families, prompting strategies, and task structures. Moreover, scalar diversity gradients emerge only in specific model-condition combinations, suggesting that pragmatic reasoning in LLMs reflects an interaction between internal probabilistic representations and task-induced prompting behavior rather than a stable competence captured by a single evaluation paradigm. These findings highlight the central role of evaluation design in interpreting pragmatic abilities in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。