arXiv:2508.04681cs.CV2025-08ICCV被引 10

构建首个第一人称视角下人-物-人交互数据集,推动智能助手物理世界应用

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

  • 基于第一人称视觉与语音指令,构建人-物-人协作任务框架
  • 产出1.2M帧、11.4小时的多模态数据,含精确动作与语音命令
  • 提供动作估计、交互生成与预测三大新基准,适合具身智能研究者

从真实世界人机交互数据中学习行动模型对构建高效通用智能助手至关重要。然而,现有数据集多聚焦特定交互类别,且忽视助手以第一人称视角感知与行动的特性。本文主张通用交互知识与第一人称模态均不可或缺。我们设计了基于人工辅助任务的视觉-语言-行动框架,助手根据第一人称视觉和指令为指导者提供服务。通过混合RGB-MoCap系统,助教与指导者在多物体场景中按GPT生成脚本互动。在此设定下,我们构建了InterVLA——首个大规模人-物-人交互数据集,包含11.4小时、120万帧的多模态数据,覆盖2个第一人称与5个第三人称视频,具备精准的人体/物体运动及语音指令。同时建立新的第一人称人体动作估计、交互合成与预测基准,并进行详尽分析。我们认为InterVLA测试平台与基准将推动物理世界中AI代理的未来发展。

原文摘要 · Abstract (English)

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-person acquisition. We urge that both the generalist interaction knowledge and egocentric modality are indispensable. In this paper, we embed the manual-assisted task into a vision-language-action framework, where the assistant provides services to the instructor following egocentric vision and commands. With our hybrid RGB-MoCap system, pairs of assistants and instructors engage with multiple objects and the scene following GPT-generated scripts. Under this setting, we accomplish InterVLA, the first large-scale human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, spanning 2 egocentric and 5 exocentric videos, accurate human/object motions and verbal commands. Furthermore, we establish novel benchmarks on egocentric human motion estimation, interaction synthesis, and interaction prediction with comprehensive analysis. We believe that our InterVLA testbed and the benchmarks will foster future works on building AI agents in the physical world.

第一人称人机交互数据集具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。