从单目视频学习灵巧操作,无需昂贵人工标注。
V2P-Manip: Learning Dexterous Manipulation from Monocular Human Videos

- 通过视频直接提取高保真、物理合理的动作轨迹。
- 在TACO和OakInk上平均成功率超75%,优于现有方法。
- 适合希望低成本获取灵巧操作数据的研究者。
实现自主灵巧操作需要大规模、类人的精确动作序列。为弥补昂贵遥操作数据的不足,从单目视频中提取兼具视觉真实性和物理合理性的轨迹,成为具身智能的重要方向。为此,我们提出V2P-Manip,一种高效框架,可直接从人类示范视频中学习灵巧操作策略。构建了包含3D资产获取、轨迹估计和灵巧策略学习的集成流程。为弥合视觉感知与物理约束之间的差距,引入两阶段精炼机制以保证空间对齐与物理一致性。在TACO和OakInk基准上的评估表明,本方法在姿态精度、非结构化环境适应性及训练效率方面显著优于先前方法。实验结果验证,该方法在多个合成操作任务中平均成功率达75%以上,并证明所提取的操作先验可适配多种灵巧手模型。
原文摘要 · Abstract (English)
Achieving autonomous robotic dexterous manipulation requires precise, human-like action sequences at scale. As a scalable supplement to costly teleoperation data, extracting trajectories with both visual fidelity and physical plausibility from monocular videos represents a promising frontier in embodied AI. To this end, we introduce V2P-Manip, an efficient framework designed to learn dexterous manipulation policies directly from human demonstration videos. We establish an efficient, integrated pipeline encompassing 3D asset acquisition, trajectory estimation, and dexterous policy learning. To bridge the gap between visual perception and physical constraints, we introduce a two-stage refinement process to enforce spatial alignment and physical consistency. Evaluations on the TACO and OakInk benchmarks demonstrate that our approach significantly outperforms previous methods in pose accuracy, adaptability to unstructured environments, and training efficiency. Ultimately, experimental results confirm an average success rate of over 75% across multiple synthetic manipulation tasks and validate the adaptability of the extracted manipulation priors across diverse dexterous hand embodiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。