用视频扩散模型重构被遮挡的自摄像头手部动作,效果远超现有方法。
ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

- 将视频扩散模型转为确定性几何编码器,单次前向传播即可还原遮挡部分
- 在ARCTIC和HOT3D上相比旧方法减少30%-40%的关节点误差
- 无需外部检测器或测试时相机参数,适合机器人操作数据生成
自摄像头视频为具身智能提供了可扩展的操作数据,但因物体严重遮挡和频繁离屏,恢复度量级3D手部轨迹仍具挑战。现有单帧与短时窗回归器在手短暂离屏时失效,而近期视频扩散模型(VDMs)依赖耗时且随机的多步采样作为像素空间渲染器。本文将VDM重用于确定性几何编码器,一次前向传播即可揭示当前观测之外的场景内容,包括被遮挡和离屏的手部。提出ACE-Ego-Hand,一种离线片段级框架,通过确定性干净潜在编码器提取特征,并由双向时空解码器还原。该方法无需外部检测器即可恢复连续双手轨迹并获得度量位置;另配置中基于射线的相机求解器亦无需测试时相机内参。在五个自摄像头基准上,ACE-Ego-Hand达到新最优性能,在遮挡密集的ARCTIC上降低MPJPE-p 30%,在HOT3D上降低40%。当包含离屏手部评估时,性能提升达46%-61%,为从日常人类视频到机器人操作数据提供可扩展路径。
原文摘要 · Abstract (English)
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when a hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce ACE-Ego-Hand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. ACE-Ego-Hand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that requires no test-time camera intrinsics. Across five egocentric benchmarks, ACE-Ego-Hand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%-61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。