构建首个从第一视角视频推断心智状态的评测基准
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
- 基于因果心智模型生成第一视角视频问答数据
- 大模型在目标推理上接近人类,信念与未来动作预测仍显著落后
- 推动具备用户心理建模能力的智能助手发展
我们提出EgoToM,一个将心智理论(ToM)评估拓展至第一视角视频领域的新型视频问答基准。利用因果心智模型,我们在Ego4D数据集上生成多选题,用于评估对摄像头佩戴者目标、信念及下一步行为的预测能力。我们研究了人类与当前最先进的多模态大语言模型(MLLMs)在这三类相互关联的推理任务上的表现。结果表明,尽管MLLMs在从第一视角视频推断目标时达到接近人类的准确率,但即便在参数量超过1000亿的最大模型上,其在推断佩戴者当下信念状态及最符合未见未来情节的行为时仍明显落后于人类。我们认为这些结果将影响未来具备用户内在心理状态建模能力的第一视角数字助手的设计。
原文摘要 · Abstract (English)
We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the ability to predict a camera wearer's goals, beliefs, and next actions. We study the performance of both humans and state of the art multimodal large language models (MLLMs) on these three interconnected inference problems. Our evaluation shows that MLLMs achieve close to human-level accuracy on inferring goals from egocentric videos. However, MLLMs (including the largest ones we tested with over 100B parameters) fall short of human performance when inferring the camera wearers' in-the-moment belief states and future actions that are most consistent with the unseen video future. We believe that our results will shape the future design of an important class of egocentric digital assistants which are equipped with a reasonable model of the user's internal mental states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。