arXiv:2507.18503cs.CV2025-07被引 3

用语义-中心视野机制预测人类找目标时的视线轨迹

Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention

  • 结合深度检测与概率语义融合生成动态注意力图
  • 在COCO-Search18上逼近真人视线序列,优于多数纯自上而下模型
  • 适合实时认知计算与机器人视觉系统开发人员参考

在目标导向的视觉任务中,人类感知同时受自上而下和自下而上线索引导,中心视野在高效引导注意方面起关键作用。本文评估了SemBA-FAST——一种专为目标存在场景下的主动视觉搜索设计的自上而下框架,用于预测人类视觉注意。该方法将深度目标检测与概率语义融合机制结合,动态生成注意力图,利用预训练检测器和人工中心视野,逐步更新自上而下知识以优化注视预测。我们在COCO-Search18基准数据集上对比了其与其它注视预测模型的表现,结果表明该方法生成的注视序列与真实人类轨迹高度吻合,显著优于基线及其他自上而下模型,在部分指标上可与依赖注视信息的模型媲美。这些发现揭示了语义-中心视野概率框架在模拟人类注意行为方面的潜力,对实时认知计算与机器人应用具有重要启示。

原文摘要 · Abstract (English)

In goal-directed visual tasks, human perception is guided by both top-down and bottom-up cues. At the same time, foveal vision plays a crucial role in directing attention efficiently. Modern research on bio-inspired computational attention models has taken advantage of advancements in deep learning by utilizing human scanpath data to achieve new state-of-the-art performance. In this work, we assess the performance of SemBA-FAST, i.e. Semantic-based Bayesian Attention for Foveal Active visual Search Tasks, a top-down framework designed for predicting human visual attention in target-present visual search. SemBA-FAST integrates deep object detection with a probabilistic semantic fusion mechanism to generate attention maps dynamically, leveraging pre-trained detectors and artificial foveation to update top-down knowledge and improve fixation prediction sequentially. We evaluate SemBA-FAST on the COCO-Search18 benchmark dataset, comparing its performance against other scanpath prediction models. Our methodology achieves fixation sequences that closely match human ground-truth scanpaths. Notably, it surpasses baseline and other top-down approaches and competes, in some cases, with scanpath-informed models. These findings provide valuable insights into the capabilities of semantic-foveal probabilistic frameworks for human-like attention modelling, with implications for real-time cognitive computing and robotics.

视觉搜索注意力建模语义注意机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。