arXiv:2409.02638cs.CV2024-09TPAMI被引 23

用视觉语言模型和运动感知Mamba预测手部轨迹,兼顾时序合理性和实时性。

MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos

  • 引入运动感知Mamba,融合佩戴者自身运动信息进行选择性扫描
  • 在5个公开数据集上达到领先水平,且推理速度满足实时需求
  • 无需显式物体可用性标注,适合虚拟现实与机器人操作场景

通过第一人称视角视频理解人类意图与动作是实现具身人工智能的重要一步。手部轨迹预测作为其中关键环节,有助于理解人体运动模式并支持扩展现实与机器人操作等下游任务。然而,在仅有第一人称视频的情况下,如何捕捉与时间因果一致的高层意图仍具挑战,尤其在相机自身运动干扰、缺乏物体可用性标签指导手部路径优化时更为突出。为此,本文提出MADiff方法,采用扩散模型预测未来手部关键点。其潜空间去噪过程由所提出的运动感知Mamba实现,将佩戴者自身运动信息融入,形成运动驱动的选择性扫描(MDSS)。为在无显式可用性标注下识别手与场景关系,引入融合视觉与语言特征的预训练基础模型以提取视频片段中的高层语义。在五个公开数据集上的实验表明,MADiff在新提出的评估指标及现有基准上均表现优异,且具备实时推理能力。代码与预训练模型将在项目页面公开:https://irmvlab.github.io/madiff.github.io。

原文摘要 · Abstract (English)

Understanding human intentions and actions through egocentric videos is important on the path to embodied artificial intelligence. As a branch of egocentric vision techniques, hand trajectory prediction plays a vital role in comprehending human motion patterns, benefiting downstream tasks in extended reality and robot manipulation. However, capturing high-level human intentions consistent with reasonable temporal causality is challenging when only egocentric videos are available. This difficulty is exacerbated under camera egomotion interference and the absence of affordance labels to explicitly guide the optimization of hand waypoint distribution. In this work, we propose a novel hand trajectory prediction method dubbed MADiff, which forecasts future hand waypoints with diffusion models. The devised denoising operation in the latent space is achieved by our proposed motion-aware Mamba, where the camera wearer's egomotion is integrated to achieve motion-driven selective scan (MDSS). To discern the relationship between hands and scenarios without explicit affordance supervision, we leverage a foundation model that fuses visual and language features to capture high-level semantics from video clips. Comprehensive experiments conducted on five public datasets with the existing and our proposed new evaluation metrics demonstrate that MADiff predicts comparably reasonable hand trajectories compared to the state-of-the-art baselines, and achieves real-time performance. We will release our code and pretrained models of MADiff at the project page: https://irmvlab.github.io/madiff.github.io.

手部轨迹扩散模型第一人称视觉运动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。