从第一人称视频学习可泛化的柔体物理模型,无需复杂优化即可预测未知物体变形。
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video

- 通过代码本压缩逆物理解,实现从单段第一人称视频生成柔体数字孪生。
- 在未见物体上实现高精度形变预测,零样本泛化能力超越基线方法。
- 适用于机器人规划,仅需人类操作视频即可构建可交互的内部世界模型。
人类通过日常互动自然理解物体物理特性,但准确预测复杂柔体动态(如弹性材料与织物)仍是计算机视觉与机器人领域的重大挑战。本文提出EgoPhys框架,利用第一人称RGB视频和通用先验,构建可泛化的柔体物理数字孪生。该方法通过将每类物体的逆物理解提炼为紧凑代码本,实现对未见物体密集弹簧刚度场的预测,且无需测试时逐弹簧优化。基于多样第一人称交互数据训练,EgoPhys在重建、未来预测及零样本泛化上均优于基线。为支持训练与评估,我们构建了覆盖多种柔体物体、场景与操作风格的第一人称交互数据集。在真实xArm6机械臂上部署表明,仅需一段人类玩耍视频初始化的数字孪生,即可作为内部世界表征,辅助柔体物体规划,验证第一人称RGB观测作为真实到仿真路径的可行性。
原文摘要 · Abstract (English)
Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformable dynamics, such as elastic materials and fabrics, remains a major challenge for computer vision and robotics. We present EgoPhys, a framework that constructs deformable physical digital twins from egocentric RGB-only video using generalizable priors. EgoPhys overcomes the limitations of existing methods to enable controllable deformable digital twin generation from egocentric videos by distilling per-object inverse-physics solutions into a compact codebook, enabling prediction of dense spring stiffness fields for unseen objects without per-spring test-time optimization. Trained with generalizable priors from diverse egocentric interactions, EgoPhys outperforms baselines in reconstruction, future prediction, and zero-shot generalization. To support training and evaluation, we curate an egocentric interaction dataset covering diverse deformable objects, scenes, and manipulation styles. We deploy EgoPhys on a real xArm6 robot, demonstrating that a digital twin initialized from a single egocentric human play video can serve as an internal world representation to aid in deformable-object planning, highlighting egocentric RGB observations as a scalable path toward real-to-sim pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。