用生成模型模拟说话人和听者的语言选择,让机器更懂语境中的微妙含义。
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

- 用大模型生成可能的表达方式,再由规则模块筛选最优解。
- 在三个语用任务中表现接近人类,尤其在生成表达上优于传统方法。
- 适合研究语言理解、认知建模或人机对话的开发者与研究人员。
语用语言使用需要对替代选项进行推理:说话人可能选择的其他表达,或听者可能考虑的其他解释。形式化和计算化的语用模型必须明确定义对话者所推理的替代集合,这通常通过人工指定实现。本文提出SAGE框架——结合认知模型的可解释性与语言模型的生成灵活性,将语用过程分解为三类模块:提议者(利用语言模型生成开放式的候选替代表达)、评估者(评估这些替代项的语义、复杂度或典型性)和选择器(执行基于认知动机的任务分析规则)。我们在三个案例研究中评估SAGE:指称表达生成、方式隐含意义(M- implicatures)和格莱斯会话隐含意义。通过计算认知建模的标准方法(如消融实验、基线对比、与人类数据的定量拟合)进行严格评估。结果显示,SAGE模型整体准确率高,常优于基线;但组件分析揭示不对称性:语言模型作为提议者能可靠生成适合语用建模的替代项,而作为评估者则更适合提供直觉判断,而非理论或形式化度量的判断。讨论了神经符号模型作为人类语用语言使用解释性范式的优势与局限。
原文摘要 · Abstract (English)
Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification. Here we propose a framework, ScAffolded Generative models for Explanation (SAGE), that combines the explanatory transparency of cognitive models with the generative flexibility of language models (LMs). SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives (e.g., their semantics, complexity, or typicality); and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis. We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner (M-)implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data. Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。