arXiv:2603.25539cs.CV2026-03被引 1

从大量第一人称视频中自动提取物体运动结构,提升机器人交互能力

PAWS: Perception of Articulation in the Wild at Scale from Egocentric Videos

  • 直接从真实场景的视频中解析物体开合动作
  • 在HD-EPIC和Arti4D数据集上显著超越基线方法
  • 可直接用于机器人抓取与3D模型微调,适合具身智能研究者

关节运动感知旨在恢复物体(如抽屉、柜门)的运动与结构,是机器人、仿真和动画领域三维场景理解的基础。现有基于学习的方法依赖高质量3D数据与人工标注进行监督训练,难以规模化且多样性不足。为此,我们提出PAWS,一种直接从大规模真实场景第一人称视频中提取手-物交互下的物体关节信息的方法。我们在公开数据集HD-EPIC与Arti4D上评估该方法,结果显著优于基线。进一步实验表明,所提取的关节信息可有效提升下游任务表现,包括微调3D关节预测模型及支持机器人操作。项目主页:https://aaltoml.github.io/PAWS/

原文摘要 · Abstract (English)

Articulation perception aims to recover the motion and structure of articulated objects (e.g., drawers and cupboards), and is fundamental to 3D scene understanding in robotics, simulation, and animation. Existing learning-based methods rely heavily on supervised training with high-quality 3D data and manual annotations, limiting scalability and diversity. To address this limitation, we propose PAWS, a method that directly extracts object articulations from hand-object interactions in large-scale in-the-wild egocentric videos. We evaluate our method on the public data sets, including HD-EPIC and Arti4D data sets, achieving significant improvements over baselines. We further demonstrate that the extracted articulations benefit downstream tasks, including fine-tuning 3D articulation prediction models and enabling robot manipulation. See the project website at https://aaltoml.github.io/PAWS/.

关节感知第一人称视频机器人操作自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。