arXiv:2409.14986cs.CLcs.AI2024-09被引 1

让大模型预测对话中他人信念的不确定性,突破传统二元思维。

Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting

  • 以对话预测为场景,让模型估计对话者对自己观点的不确定程度。
  • 模型最多能解释7%的他人信念不确定性方差,表现有限。
  • 适合研究心理建模、人机交互与社会推理的学者参考。

传统心智理论通常将他人信念视为非黑即白,但若对方对自己的信念本身存疑呢?如何量化这种不确定性?本文提出一套新任务,要求语言模型在对话中预测他人信念的不确定性(以概率形式表示)。任务基于对话预测设计,将对话参与者视为预测主体,要求模型估计其信念的不确定程度。我们在三个对话语料库(社交、谈判、任务导向)上,对八种语言模型进行了实验,考察了重标度方法、方差减少策略及人口统计背景的影响。结果显示,模型最多可解释7%的他人信念不确定性方差,表明该任务极具挑战性,未来在实际应用(如识别错误信念)中仍有巨大提升空间。

原文摘要 · Abstract (English)

Typically, when evaluating Theory of Mind, we consider the beliefs of others to be binary: held or not held. But what if someone is unsure about their own beliefs? How can we quantify this uncertainty? We propose a new suite of tasks, challenging language models (LMs) to model the uncertainty of others in dialogue. We design these tasks around conversation forecasting, wherein an agent forecasts an unobserved outcome to a conversation. Uniquely, we view interlocutors themselves as forecasters, asking an LM to predict the uncertainty of the interlocutors (a probability). We experiment with re-scaling methods, variance reduction strategies, and demographic context, for this regression task, conducting experiments on three dialogue corpora (social, negotiation, task-oriented) with eight LMs. While LMs can explain up to 7% variance in the uncertainty of others, we highlight the difficulty of the tasks and room for future work, especially in practical applications, like anticipating ``false

心智理论对话预测不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。