arXiv:2609.06208cs.CV2026-09

无需训练的AI agent,用眼神线索统一理解视觉目标与语义解释

From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation

论文配图:From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation
图 1 · 摘自论文原文
  • 基于预训练模型+视觉提示,不需训练即可推理眼神关注点
  • 在GazeFollow和GazeHOI上达到当前最佳性能
  • 能解释错误标签、无词汇限制,适合跨任务视觉理解

理解人类注意力对场景解析至关重要,但现有方法多依赖大量训练的模型,缺乏可解释性。以往方法难以在无大量监督下联合推理眼神目标、注意对象与视觉定位。据我们所知,本文首次提出无需训练的注视目标代理(GTA),可在注视目标预测、注意力定位和物体识别等任务中实现统一推理。该方法利用预训练视觉语言模型,结合视觉引导提示,并采用基于记忆的检索策略处理高不确定性样本,从而提升性能而无需额外训练。我们在定量和定性两方面评估该方法:定量上,在GazeFollow和GazeHOI基准上达到最先进性能;定性上,能提供详细语义预测,即使标注错误也能正确识别目标,且不受词汇表限制。

原文摘要 · Abstract (English)

Understanding human attention is fundamental for scene interpretation, yet existing approaches often rely on heavily trained models that lack interpretability. Prior methods struggle to jointly reason about gaze targets, attended objects, and visual grounding without extensive supervision. To the best of our knowledge, this work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification. This is achieved by leveraging pretrained vision-language models, augmenting them with visually guided prompts, and employing a memory-based retrieval strategy for high-uncertainty samples to improve performance without additional training. We evaluate our approach using both quantitative metrics and qualitative results. Quantitatively, our method achieves state of the art performance on the GazeFollow and GazeHOI benchmarks. Qualitatively, our agent provides detailed semantic predictions, predicts the correct targets even when ground truth labels are wrong, and remains flexible without vocabulary constraints.

视觉理解注视推理零样本多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。