arXiv:2505.18497cs.CL2025-05Conference of the …被引 3

通过对比推理测试,发现大模型的语用能力随训练逐步增强。

The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models

  • 设计对比性语用数据集,探测模型对说话人意图的推断能力
  • 基线模型已具语用敏感性,微调后能力显著提升
  • 适合关注模型沟通能力与人类语言规范对齐的研究者

当前大语言模型在隐含意义解析和心智理论推理等社交智能任务中展现出初步能力,这些均需较强的语用理解。然而,模型如何在训练过程中习得这种语用能力仍不明确。本文提出ALTPRAG,一个基于语用中‘替代选项’概念的数据集,用于评估不同训练阶段的22个LLM是否能准确推断细微的说话人意图。每个样本包含两个同样合理但语用上不同的后续语句,要求模型(1)推断说话人的真实意图,(2)解释为何说话人会选择某一表达而非其替代项,从而通过对比推理直接探测语用能力。系统评估覆盖预训练、监督微调(SFT)和偏好优化三个关键阶段,结果表明:即使基础模型也表现出明显的语用敏感性,且随着模型规模和数据量增加持续提升;SFT与强化学习人类反馈(RLHF)进一步推动认知-语用场景下的性能增长。研究揭示语用能力是模型训练中的涌现性与组合性特征,为模型与人类交际规范对齐提供了新洞见。

原文摘要 · Abstract (English)

Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning, both of which require substantial pragmatic understanding. However, how LLMs acquire this pragmatic competence throughout the training process remains poorly understood. In this work, we introduce ALTPRAG, a dataset grounded in the pragmatic concept of alternatives, to evaluate whether LLMs at different training stages can accurately infer nuanced speaker intentions. Each instance pairs two equally plausible yet pragmatically divergent continuations and requires the model to (i) infer the speaker's intended meaning and (ii) explain when and why a speaker would choose one utterance over its alternative, thus directly probing pragmatic competence through contrastive reasoning. We systematically evaluate 22 LLMs across 3 key training stages: after pre-training, supervised fine-tuning (SFT), and preference optimization, to examine the development of pragmatic competence. Our results show that even base models exhibit notable sensitivity to pragmatic cues, which improves consistently with increases in model and data scale. Additionally, SFT and RLHF contribute further gains, particularly in cognitive-pragmatic scenarios. These findings highlight pragmatic competence as an emergent and compositional property of LLM training and offer new insights for aligning models with human communicative norms.

语用理解大模型能力对话推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。