arXiv:2604.02512cs.CLcs.AI2026-04

LLM社会推理结构像人,但强度不准,用语用理论提示可改善。

Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting

论文配图:Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
图 1 · 摘自论文原文
  • 用效应量比和校准偏差分评估模型社会推理的结构与强度匹配度。
  • 三款前沿模型均复现人类社会推断结构,但强度偏差显著,尤其在数值模糊性任务中。
  • 结合说话者知识与意图提示能最有效降低偏差,是当前最优改进策略。

大型语言模型(LLMs)日益表现出类人的语用与社会推理模式。本文聚焦两个问题:LLMs是否不仅在定性上,也在定量上逼近人类社会意义;基于语用理论的提示策略能否提升这种逼近?为此,我们提出两个校准导向指标:效应量比(ESR)和校准偏差分(CDS),用于区分结构保真度与幅度校准。基于两项语用假设——社会意义源于对语言替代项的推理,以及听者会推断说话者知识状态与交际动机——我们设计提示策略。在三个前沿模型上的数值(不)精确性案例研究显示,所有模型可靠复现了人类社会推断的定性结构,但在幅度校准上差异显著。提示模型推理说话者知识与动机,最一致地减少幅度偏差;而仅提示替代项意识则加剧夸大。两者结合是唯一能在所有模型中同时提升各项校准敏感指标的干预,尽管精细幅度校准仍部分未解。因此,LLMs捕捉到推断结构,但对推断强度的扭曲程度各异,语用理论提供了有用但不完整的优化路径。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related questions: do LLMs approximate human social meaning not only qualitatively but also quantitatively, and can prompting strategies informed by pragmatic theory improve this approximation? To address the first, we introduce two calibration-focused metrics distinguishing structural fidelity from magnitude calibration: the Effect Size Ratio (ESR) and the Calibration Deviation Score (CDS). To address the second, we derive prompting conditions from two pragmatic assumptions: that social meaning arises from reasoning over linguistic alternatives, and that listeners infer speaker knowledge states and communicative motives. Applied to a case study on numerical (im)precision across three frontier LLMs, we find that all models reliably reproduce the qualitative structure of human social inferences but differ substantially in magnitude calibration. Prompting models to reason about speaker knowledge and motives most consistently reduces magnitude deviation, while prompting for alternative-awareness tends to amplify exaggeration. Combining both components is the only intervention that improves all calibration-sensitive metrics across all models, though fine-grained magnitude calibration remains only partially resolved. LLMs thus capture inferential structure while variably distorting inferential strength, and pragmatic theory provides a useful but incomplete handle for improving that approximation.

大模型语用学社会推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。