arXiv:2505.16602cs.CV2025-05NeurIPS被引 12

从第一视角视频和文本生成逼真手物交互动作,解决视角不稳与泛化难问题。

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

  • 分层设计:高层用视觉语言模型推断动作先验,底层用扩散模型生成精细轨迹
  • 在5个域内和2个跨域数据集上,手腕位移误差降低86.9%,关节旋转误差降34.1%
  • 适配新物体和复杂场景,适合虚拟现实、机器人模仿等应用

第一视角手物交互动作生成对沉浸式增强/虚拟现实和机器人模仿至关重要,但受视角不稳、自遮挡、透视畸变和噪声自我运动影响。现有方法依赖预设3D物体先验,限制了对新物体的泛化能力;而近期多模态方法存在文本提示模糊、3D手物关联建模复杂及开环预测误差累积等问题。我们提出MEgoHand,一种从第一视角RGB图像、文本和初始手部姿态生成物理合理手物交互的动作框架。其采用双层级架构:高层‘大脑’利用视觉语言模型(VLM)结合单目深度估计器,实现无物体依赖的空间推理;低层基于DiT的流匹配策略,通过时间正交滤波提升轨迹稳定性。为解决数据不一致问题,设计逆MANO重定向网络与虚拟RGB-D渲染器,构建统一数据集,包含335万帧RGB-D图像、2.4万次交互和1200个物体。大量实验在五个域内和两个跨域数据集上验证有效性,腕部平移误差减少86.9%,关节旋转误差降低34.1%,展现对细粒度手部结构的精准建模能力和跨场景鲁棒泛化能力。

原文摘要 · Abstract (English)

Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, self-occlusions, perspective distortion, and noisy ego-motion. Existing methods rely on predefined 3D object priors, limiting generalization to novel objects, which restricts their generalizability to novel objects. Meanwhile, recent multimodal approaches suffer from ambiguous generation from abstract textual cues, intricate pipelines for modeling 3D hand-object correlation, and compounding errors in open-loop prediction. We propose MEgoHand, a multimodal framework that synthesizes physically plausible hand-object interactions from egocentric RGB, text, and initial hand pose. MEgoHand introduces a bi-level architecture: a high-level "cerebrum" leverages a vision language model (VLM) to infer motion priors from visual-textual context and a monocular depth estimator for object-agnostic spatial reasoning, while a low-level DiT-based flow-matching policy generates fine-grained trajectories with temporal orthogonal filtering to enhance stability. To address dataset inconsistency, we design a dataset curation paradigm with an Inverse MANO Retargeting Network and Virtual RGB-D Renderer, curating a unified dataset of 3.35M RGB-D frames, 24K interactions, and 1.2K objects. Extensive experiments across five in-domain and two cross-domain datasets demonstrate the effectiveness of MEgoHand, achieving substantial reductions in wrist translation error (86.9%) and joint rotation error (34.1%), highlighting its capacity to accurately model fine-grained hand joint structures and generalize robustly across diverse scenarios.

动作生成多模态第一视角手物交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。