arXiv:2604.20361cs.CV2026-04

用语言描述找物体时,模型能更准预测人眼注视路径。

Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models

论文配图:Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
图 1 · 摘自论文原文
  • 用视觉语言模型融合图像与语言特征,提升感知对齐
  • 引入历史注视位置信息,使预测路径更合理
  • 轻量级辅助模块提升定位精度,不增加计算负担

目标指代引导的注视路径预测(ORSP)旨在根据语言描述预测人在视觉场景中寻找特定目标物体时的人眼注视轨迹。多模态信息融合是ORSP的关键。为此,我们提出新模型ScanVLA,首先利用视觉语言模型(VLM)从输入图像和指代表达中提取并融合内在对齐的视觉与语言特征表示。为进一步增强模型对细粒度位置信息的感知,我们不仅设计了历史增强注视解码器(HESD),直接以历史注视位置作为输入,帮助预测当前注视点更合理的位置;还采用冻结的分割LoRA作为辅助组件,更精准定位被指代物体,从而在不显著增加计算和时间开销的前提下提升注视路径预测性能。大量实验表明,ScanVLA在目标指代引导的注视路径预测任务中显著优于现有方法。

原文摘要 · Abstract (English)

Object Referring-guided Scanpath Prediction (ORSP) aims to predict the human attention scanpath when they search for a specific target object in a visual scene according to a linguistic description describing the object. Multimodal information fusion is a key point of ORSP. Therefore, we propose a novel model, ScanVLA, to first exploit a Vision-Language Model (VLM) to extract and fuse inherently aligned visual and linguistic feature representations from the input image and referring expression. Next, to enhance the ScanVLA's perception of fine-grained positional information, we not only propose a novel History Enhanced Scanpath Decoder (HESD) that directly takes historical fixations' position information as input to help predict a more reasonable position for the current fixation, but also adopt a frozen Segmentation LoRA as an auxiliary component to help localize the referred object more precisely, which improves the scanpath prediction task without incurring additional large computational and time costs. Extensive experimental results demonstrate that ScanVLA can significantly outperform existing scanpath prediction methods under object referring.

视觉语言模型注视预测多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。