arXiv:2609.08607cs.CV2026-09

基于双目视觉的交互场估计新方法,精度达27.96毫米

GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation

论文配图:GOLF: Global Observation with Local Focus for Calibration-Aware Stereo Interaction Field Estimation
图 1 · 摘自论文原文
  • 融合全局上下文与局部手物证据,结合普鲁克射线几何建模
  • 在隐藏测试集上实现27.96毫米平均轨迹误差,官方得分27.61
  • 适合需要高精度手物交互理解的机器人操作场景

我们提出GOLF,这是在HANDS@ECCV 2026举办的SHOW3D交互场估计挑战赛中的第一名解决方案。给定同步的头戴式双目视觉图像,任务是预测21个手关节到所操作物体最近点的3D向量。GOLF结合密集全局上下文、局部采样的手/物体证据以及公共帧下的Plücker射线几何。我们采用DINOv3 ViT-H+/16并引入LoRA与可训练的LayerNorm参数,联合解码两个交互场。主模型在隐藏测试集上取得官方得分27.61,平均ADE为27.96毫米;与一个互补的直接微调变体进行等权集成后,官方得分提升至27.47,平均ADE降至27.82毫米,获得第一名。

原文摘要 · Abstract (English)

We present GOLF, the first-place solution to the SHOW3D Interaction Field Estimation Challenge at HANDS@ECCV 2026. Given synchronized egocentric stereo views, the task is to predict a 3D vector from each of 21 hand joints to the closest point on the manipulated object. GOLF combines dense global context, locally sampled hand/object evidence, and common-frame Pl\"ucker-ray geometry. We adapt DINOv3 ViT-H+/16 with LoRA and trainable LayerNorm parameters, then jointly decode both interaction fields. Our primary model achieves an official score of 27.61 and a mean ADE of 27.96 mm on the hidden test set. An equal-weight ensemble with a complementary directly fine-tuned variant improves these results to an official score of 27.47 and a mean ADE of 27.82 mm, securing first place.

立体匹配交互估计手物交互视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。