arXiv:2608.06192cs.CV2026-08

从单目视频中预测手部与物体接触的实时压力分布。

HOPE: Hand-Object Pressure Estimation from Monocular Videos

论文配图:HOPE: Hand-Object Pressure Estimation from Monocular Videos
图 1 · 摘自论文原文
  • 将压力估计转化为以手为中心的视频预测任务,输出与物体形状无关。
  • 在多个数据集上实现跨场景、跨传感器的压力与接触联合预测。
  • 适用于裸手、日常场景,超越传统平面压力或仅接触检测方法。

从视觉中估计物理压力对于理解丰富的手物交互至关重要。然而,以往基于视觉的压力估计方法大多局限于平面表面和单图输入,难以应用于具有多样化物体的动态手物交互。本文提出一种新范式:将压力估计建模为以手为中心的单目视频预测问题,直接在手网格顶点上预测随时间演化的法向压力和接触信息,输出空间独立于物体形状和传感器布局。在此基础上,提出HOPE框架,包含两个核心组件:首先,将触觉手套压力、平面传感器压力及基于距离的手物接触标注统一映射到共享的手顶点空间,使无度量标签的裸手数据可辅助压力学习;其次,引入顶点锚定视频变换器,将每个顶点视为持续存在的标记,聚合时序视觉特征与手姿态,并采用接触门控压力头,确保无接触时压力为零。在OpenTouch、PressureVisionDB及手物交互基准测试中验证,尽管主要依赖戴手套视频的度量压力监督,HOPE仍能泛化至裸手自摄像和真实场景视频,实现超出接触仅检测或平面压力基线的联合接触与压力预测。

原文摘要 · Abstract (English)

Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects. We instead formulate pressure estimation as a hand-centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per-vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose \textbf{HOPE}, a framework with two key components. First, we lift tactile-glove pressure, planar-sensor pressure, and distance-based hand-object contact annotations into a shared hand vertex space, allowing bare-hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex-anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact-gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand-object contact benchmarks validate HOPE across object-pressure, surface-pressure, and contact-supervised HOI settings. Despite using metric pressure supervision primarily from gloved-hand videos, HOPE generalizes to bare-hand egocentric and in-the-wild videos, producing joint contact and pressure predictions beyond the scope of contact-only or planar-pressure baselines.

压力估计手物交互视频预测单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。