构建首个聚焦社交推断的对话数据集,评测大模型理解讽刺等社会隐含意义的能力。
SocialNLI: A Dialogue-Centric Social Inference Dataset
- 设计包含讽刺、反语等复杂社交细节的对话数据集
- 通过多步反事实推理测试模型社会认知能力,发现现有模型表现不佳
- 适合研究社会智能、具身认知与对话理解的AI研究人员
从人类对话中进行心理理论推断是衡量模型潜在社交能力的重要指标,而这一能力对智能助手至关重要。然而,当前的大语言模型和推理模型在理解对话数据中的复杂社交现象(如讽刺和反语)时仍存在显著困难。为评估现有模型的缺陷并探索解决方案,我们提出 SocialNLI(SoNLI)——首个专注于社交对话推断的数据集。该数据集包含精心挑选的对话转录文本,聚焦于讽刺、反语等复杂社交细节,并配以推断判断、置信度评分及人工撰写解释。我们以心理理论为视角,通过多步反事实推理评估大模型的社会认知能力,揭示其在理解深层社会语境方面的不足。
原文摘要 · Abstract (English)
Making theory-of-mind inferences from human dialogue is a strong indicator of a model's underlying social abilities, which are fundamental for adept AI assistants. However, large language and reasoning models struggle to understand sophisticated social phenomena in transcript data, such as sarcasm and irony. To assess the weaknesses of current models and to identify their solutions, we introduce SocialNLI (SoNLI) -- the first social dialogue inference dataset. SoNLI consists of a collection of dialogue transcripts hand-picked to center complex social nuances like irony and sarcasm, paired with inferences, corresponding likelihood scores, and human-written explanations. We explore social inference analysis as a facet of theory-of-mind, and evaluate LLM and reasoning model theory-of-mind ability through multi-step counterfactual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。