arXiv:2503.23300cs.CVcs.RO2025-03被引 3

用视觉和身体数据预测人的动作,提升机器人交互能力

Learning Predictive Visuomotor Coordination

  • 构建多模态时序表征,融合第一视角视觉与身体运动信号
  • 在EgoExo4D数据集上实现高精度、连贯的动作预测
  • 适合研究人机交互与行为建模的科研人员参考

理解并预测人类的视觉-运动协调对机器人、人机交互和辅助技术至关重要。本文提出一种基于预测的任务来建模视觉-运动协调,目标是从第一人称视觉和运动学观测中预测头部姿态、视线方向及上半身运动。我们提出一种视觉-运动协调表征(Visuomotor Coordination Representation, VCR),学习多模态信号间的结构化时序依赖关系。扩展了一种基于扩散模型的运动建模框架,整合第一人称视觉与运动序列,实现时间上一致且准确的视觉-运动预测。方法在大规模EgoExo4D数据集上评估,展现出在多样化真实活动中的强泛化能力。结果表明,多模态融合对理解视觉-运动协调具有重要意义,推动了视觉-运动学习与人类行为建模的研究。

原文摘要 · Abstract (English)

Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive technologies. This work introduces a forecasting-based task for visuomotor modeling, where the goal is to predict head pose, gaze, and upper-body motion from egocentric visual and kinematic observations. We propose a \textit{Visuomotor Coordination Representation} (VCR) that learns structured temporal dependencies across these multimodal signals. We extend a diffusion-based motion modeling framework that integrates egocentric vision and kinematic sequences, enabling temporally coherent and accurate visuomotor predictions. Our approach is evaluated on the large-scale EgoExo4D dataset, demonstrating strong generalization across diverse real-world activities. Our results highlight the importance of multimodal integration in understanding visuomotor coordination, contributing to research in visuomotor learning and human behavior modeling. Project Page: https://vjwq.github.io/VCR/.

视觉运动多模态扩散模型行为预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。