arXiv:2606.06872cs.CVcs.AI2026-06中稿 · IEEE ICASSP 2026被引 1

用视觉估计第一人称视角的手压分布,更准更连贯。

EgoPressDiff: Multimodal Video Diffusion for Egocentric UV-Domain Hand-Pressure Estimation

论文配图:EgoPressDiff: Multimodal Video Diffusion for Egocentric UV-Domain Hand-Pressure Estimation
图 1 · 摘自论文原文
  • 融合手姿、3D网格和深度信息,生成物理合理的压力图。
  • 在EgoPressure数据集上,体积交并比提升超34%,误差更低。
  • 适合做增强现实、机器人模仿与人体工学分析的研究者。

从第一人称视角估计手与表面接触压力对增强现实/虚拟现实设备、机器人模仿与人体工学分析至关重要。现有方法常将压力信号离散化且独立处理帧,导致量化误差和时间不一致。本文提出EgoPressDiff,一种条件视频扩散框架,可从视觉输入生成UV压力图。核心是多模态条件策略,引入PoseNet与顶点编码器,高效提取手姿与3D网格顶点特征。这些信号结合深度信息,引导生成过程以确保压力场具有物理合理性。为有效融合异构特征,进一步提出分布校准空间层,对齐其统计特性后再融合。在EgoPressure第一人称设置下评估,EgoPressDiff达到当前最优性能,相对基线体积交并比提升超过34%,同时降低平均绝对误差并保持高时间一致性。

原文摘要 · Abstract (English)

Estimating hand-surface contact pressure from an egocentric view is crucial for AR/VR devices, robotic imitation, and ergonomic analysis. Existing methods often discretize pressure signal and process frames independently, leading to quantization errors and temporal inconsistencies. We present \emph{EgoPressDiff}, a conditional video diffusion framework that generates UV-pressure maps from visual input. The core of our approach is a multi-modal conditioning strategy, introducing a PoseNet and a Vertex Encoder to efficiently extract features from hand pose and 3D mesh vertices. These signals, along with depth information, guide the generative process to ensure the pressure fields are physically grounded. To effectively fuse these heterogeneous features, we further propose a Distribution-Calibrated Spatial Layer, which aligns their statistical properties before combination. Evaluated on the EgoPressure ego-view setting, EgoPressDiff achieves state-of-the-art results, improving Volumetric IoU by over 34\% relative to prior baseline, while reducing MAE and maintaining high temporal accuracy. Our project page is at https://egopressdiff.github.io/.

视频生成扩散模型压力估计第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。