arXiv:2607.06691cs.CV2026-07中稿 · ECCV

构建多视角协作数据集,让AI理解人类合作中的心理状态与互动。

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

论文配图:CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
图 1 · 摘自论文原文
  • 采集烹饪场景下多人多视角视频、眼动与3D环境数据
  • 建立联合注意、物体交互预测等三大基准任务
  • 适合研究社交认知、人机协作与主动辅助系统的学者

人类协作是日常生活和专业团队工作中不可或缺的部分。尽管已有大量研究关注协调与任务执行,但支持协作的认知过程(如心智理论)在真实场景中仍难以研究。为此,我们提出一个新的以自我为中心和外部视角相结合的视频数据集,涵盖真实烹饪场景中的多人协作。该数据集整合了多视角视频、高质量音频、眼动追踪和3D场景与物体扫描,并标注了共享注意力、社会线索与交互、以及代理与物体间的互动。我们建立了联合注意估计、社会条件下的物体交互预测和协作交接预测三个基准任务,推动多模态感知、主动协助与协作规划的研究。通过提供时间对齐、丰富标注的多模态数据,CoMind有助于开发与评估能够建模复杂社交互动并推理人类协作行为的AI系统。数据集与基准已公开于https://comind.ethz.ch/。

原文摘要 · Abstract (English)

Human-human collaboration is a fundamental aspect of everyday life, essential to success in a wide range of goal-directed activities from household tasks to professional teamwork. While much research has focused on modeling coordination and task execution, the cognitive processes that support such collaboration, particularly Theory of Mind (the ability to infer the mental states of others), remain difficult to study in natural settings. To address this gap, we introduce a novel egocentric and exocentric video dataset capturing real-world collaboration in cooking scenarios. The dataset integrates multi-perspective video, high-quality audio, gaze tracking, and 3D scene and object scans, with annotations for shared attention to objects, social cues and interactions between agents, as well as agent-object interactions. We establish benchmarks for Joint Attention Estimation, Socially Conditioned Object Interaction Anticipation, and Collaborative Handover Prediction, enabling research on multimodal perception, proactive assistance, and collaborative planning. By providing temporally aligned, richly annotated multimodal data, CoMind facilitates the development and evaluation of AI systems capable of modeling complex social interactions and reasoning about human behaviors in collaborative environments. Our dataset and benchmarks are made available at https://comind.ethz.ch/.

人机协作多模态感知心智理论视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。