从单目图像中同时推断手物接触点与三维受力分布,实现物理交互理解。
EgoPHI: Estimating Contact and Force from Egocentric Vision

- 基于物理仿真生成稠密顶点级力标注,解决真实力数据稀缺问题。
- 在模拟、分布外及真实场景中均优于现有方法,受力估计误差降低18%。
- 适用于需要物理感知的智能助手、机器人抓取等应用。
从第一人称视觉理解手物交互对建模人类与环境的物理互动至关重要。然而,仅定位接触点不足以刻画真实物理交互,需进一步估计作用于手部和物体上的力。本文提出EgoPHI,首个从单张单目RGB图像和物体几何结构中联合估计手与物体网格上密集接触图与三维力分布的方法。为解决大规模真实力标注缺失问题,我们设计了基于物理的仿真流程,将现有手物数据集扩展为包含逐顶点力监督的增强数据。EgoPHI学习在交互手与可动物体网格上进行稠密3D接触与力估计,突破了传统图像空间或平面设定的限制。在分布内与分布外基准测试中,其力估计性能优于现有方法,并能泛化至未见数据集。为进一步评估仿真到真实的迁移能力,我们构建了两个可捕捉密集接触与力大小的实物对象,记录了八名参与者在多种触碰与抓握类型下的交互数据。结果表明,EgoPHI在模拟、分布外及真实场景中均能恢复有意义的3D接触与力分布,推动第一人称手物理解从接触定位迈向物理驱动的交互推理。
原文摘要 · Abstract (English)
Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。