arXiv:2607.18827cs.CV2026-07中稿 · ACM Multimedia 202…

提出新基准与方法,让模型能识别未知物体的注视目标。

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

论文配图:Open-Vocabulary Gaze Object Prediction: Benchmark and Method
图 1 · 摘自论文原文
  • 用文本驱动发现候选物体,再通过凝视引导筛选目标
  • 在86类真实场景物体上实现开放词汇预测,性能超越旧方法
  • 适合关注人机交互、视觉理解的科研与工程人员

注视目标预测(GOP)旨在定位并识别人类关注的物体,对理解以人为中心的交互至关重要。然而,现有方法多基于封闭词汇范式,使用固定标签空间和特定场景数据集进行训练与评估,难以应对真实场景中注视目标呈现长尾分布或属于未见类别的问题。为此,我们构建了包含86个真实场景类别的新基准DiSG,支持开放词汇GOP(OVGOP)评估。基于DiSG,我们提出一种框架:利用文本驱动的物体发现定位潜在注视候选,再通过凝视引导选择模块精确定位实际目标。此外,为更好捕捉多样真实类别间的语义知识,引入梯度感知选择性微调(GIST),仅更新与当前类别词汇最相关的参数。大量实验表明,所提模型在开放词汇设置下表现优异,同时在传统封闭词汇设置中也优于现有方法。基准与代码已开源。

原文摘要 · Abstract (English)

Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. However, existing methods are typically trained under a closed-vocabulary paradigm with a fixed label space and evaluated on scene-specific datasets, limiting their applicability to real-world scenarios where gaze targets often follow a long-tail distribution or belong to unseen categories. To address this gap, we introduce Diverse Scenes for Gaze object prediction (DiSG), a benchmark containing 86 in-the-wild categories that facilitates the evaluation of Open-Vocabulary GOP (OVGOP). Building on DiSG, we propose a framework that leverages text-driven object discovery to localize potential gaze candidates, with a gaze-guided selection module to pinpoint the intended target from the candidate objects. Furthermore, to better capture semantic knowledge across diverse in-the-wild categories, we introduce Gradient-Informed Selection Tuning (GIST) to selectively update parameters most relevant to a given class vocabulary. Extensive experiments demonstrate that our proposed model performs effectively in open-vocabulary settings and also outperforms existing methods in the conventional closed-vocabulary setting. The benchmark and code is available at https://github.com/sensniu/ovgop.

注视预测开放词汇视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。