构建动态对话数据集,评估模型对多方信息与心理状态的推理能力
$\texttt{DIAMONDs}$: A Dataset for $\mathbb{D}$ynamic $\mathbb{I}$nformation $\mathbb{A}$nd $\mathbb{M}$ental modeling $\mathbb{O}$f $\mathbb{N}$umeric $\mathbb{D}$iscussions
- 设计可扩展方法生成高质量对话问答对,聚焦数值信息追踪
- 模型在处理虚假信念和信息不足场景时表现不佳,暴露出心智理论短板
- 适合研究多角色对话、认知推理与语言模型局限性的学者使用
理解多方对话需要强大的心理理论(ToM)能力,包括动态信息跟踪、知识不对称管理以及在长对话中区分相关信息。为推动此类场景下的ToM评估,我们提出一种精心设计的可扩展生成方法,创建了名为$ exttt{DIAMONDs}$的新对话问答数据集,涵盖常见商务、金融等群体互动。在这些目标导向的对话中,参与者需追踪特定数值量(如预期利润),其值由其他变量(如营销支出、预期销售额、薪资等)推导而来,且随对话进程变化。$ exttt{DIAMONDs}$的问题围绕这些关注量提出简单的数值推理任务(如慈善活动所需资金、下季度公司预期利润等),可在对话上下文中精确评估模型对参与者知识状态的跟踪与推理能力。对前沿语言模型的评估显示,模型在处理以参与者为中心的推理时存在显著困难,尤其在涉及虚假信念的情境下;同时对含干扰项的对话表现差,且难以识别信息不足场景。这些发现揭示了当前模型在真实多角色对话中心理理论能力的局限。
原文摘要 · Abstract (English)
Understanding multiparty conversations demands robust Theory of Mind (ToM) capabilities, including the ability to track dynamic information, manage knowledge asymmetries, and distinguish relevant information across extended exchanges. To advance ToM evaluation in such settings, we present a carefully designed scalable methodology for generating high-quality benchmark conversation-question pairs with these characteristics. Using this methodology, we create $\texttt{DIAMONDs}$, a new conversational QA dataset covering common business, financial or other group interactions. In these goal-oriented conversations, participants often have to track certain numerical quantities (say $\textit{expected profit}$) of interest that can be derived from other variable quantities (like $\textit{marketing expenses, expected sales, salary}$, etc.), whose values also change over the course of the conversation. $\texttt{DIAMONDs}$ questions pose simple numerical reasoning problems over such quantities of interest (e.g., $\textit{funds required for charity events, expected company profit next quarter}$, etc.) in the context of the information exchanged in conversations. This allows for precisely evaluating ToM capabilities for carefully tracking and reasoning over participants' knowledge states. Our evaluation of state-of-the-art language models reveals significant challenges in handling participant-centric reasoning, specifically in situations where participants have false beliefs. Models also struggle with conversations containing distractors and show limited ability to identify scenarios with insufficient information. These findings highlight current models' ToM limitations in handling real-world multi-party conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。