arXiv:2507.21888cs.CV2025-07中稿 · WACV 2026

用双方向热图融合提升指物理解的准确率

CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding

  • 提出头到指尖与腕到指尖双路径热图表示
  • 在YouRefIt上达75.0 mAP,优于现有方法
  • 适合需要精准指物理解的机器人场景

我们解决具身指物理解任务,即根据手势和语言预测说话人指向的物体。该任务需融合文本、视觉指针线索与场景上下文,但现有方法常未能充分利用视觉消歧信号。我们观察到,指代物虽常与头到指尖方向一致,但更多情况下更接近腕到指尖方向,单一方向假设过于局限。为此,我们提出双模型框架:一个基于头到指尖方向,另一个基于腕到指尖方向。引入高斯射线热图表示这两条线,并作为强监督信号引导模型关注指针线索。为融合二者互补优势,设计基于CLIP特征的指针集成模块。进一步加入辅助物体中心预测头以增强定位。在YouRefIt上取得75.0 mAP(IoU=0.25),并实现最优的CLIP与C_D得分;在未见数据集CAESAR与ISL Pointing上也表现稳健,验证了方法泛化能力。

原文摘要 · Abstract (English)

We address Embodied Reference Understanding, the task of predicting the object a person in the scene refers to through pointing gesture and language. This requires multimodal reasoning over text, visual pointing cues, and scene context, yet existing methods often fail to fully exploit visual disambiguation signals. We also observe that while the referent often aligns with the head-to-fingertip direction, in many cases it aligns more closely with the wrist-to-fingertip direction, making a single-line assumption overly limiting. To address this, we propose a dual-model framework, where one model learns from the head-to-fingertip direction and the other from the wrist-to-fingertip direction. We introduce a Gaussian ray heatmap representation of these lines and use them as input to provide a strong supervisory signal that encourages the model to better attend to pointing cues. To fuse their complementary strengths, we present the CLIP-Aware Pointing Ensemble module, which performs a hybrid ensemble guided by CLIP features. We further incorporate an auxiliary object center prediction head to enhance referent localization. We validate our approach on YouRefIt, achieving 75.0 mAP at 0.25 IoU, alongside state-of-the-art CLIP and C_D scores, and demonstrate its generality on unseen CAESAR and ISL Pointing, showing robust performance across benchmarks.

指物理解多模态视觉定位机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。