通过联合条件推理,精准定位手部关节与物体表面的交互点。
Joint-Conditioned Stereo Surface Reasoning for Interaction Field Estimation

- 以关节为条件,联合估计3D关节与表面端点位置。
- 在SHOW3D挑战赛中取得第三名,准确率显著提升。
- 适合做手物交互理解、机器人抓取等场景的应用。
预测手-物体交互场需在小而部分遮挡的图像区域中定位每个手部关节最近的物体表面点。本文将该任务视为关节条件下的表面端点估计:每个关节有其对应的最近端点,而同一手的端点可共享局部表面信息。为此提出联合条件立体表面推理(JSSR)方法:基于时序-立体网络联合预测3D关节、直接交互场及每视角端点证据。通过校准候选搜索,利用关节特异性图像匹配与跨视角对应关系评估端点假设;引入手共享候选支持机制,使关节能共享共同表面证据;并设计学习型残差门控,在观测模糊时动态控制几何修正。该系统在SHOW3D交互场挑战赛榜单中位列第三。
原文摘要 · Abstract (English)
Predicting hand--object interaction fields requires locating the nearest object-surface point for each hand joint, often from small and partially occluded image regions. We view this task as joint-conditioned surface-endpoint estimation: each joint has its own nearest endpoint, while endpoints from the same hand can draw on shared local surface evidence. This structure motivates Joint-Conditioned Stereo Surface Reasoning (JSSR). A temporal-stereo network jointly predicts 3D joints, a direct interaction field, and per-view endpoint evidence. Calibrated candidate search evaluates endpoint hypotheses using joint-specific image compatibility and cross-view correspondence. A hand-shared candidate support lets joints draw on common surface evidence, and a learned residual gate controls the geometric correction when observations are ambiguous. Our system built on this method ranked third on the SHOW3D Interaction Field Challenge leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。