arXiv:2506.02764cs.CVcs.AI2025-06中稿 · the 2025 IEEE Inte…被引 1

共享注意力表征让自由观看模型高效迁移到视觉搜索任务

Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations

  • 构建共享表征的神经网络,统一建模自由观看与视觉搜索
  • 迁移后预测注视路径误差仅增3.86%,计算量降低92.29%
  • 适合需要低延迟、少参数注意力模型的研究者

自由观看与任务驱动的视觉搜索中的人类注意力建模通常独立研究,缺乏对是否存在共用表征的探索。本文基于人类注意力变换器(HAT)提出新架构,验证该假设。结果表明,自由观看与视觉搜索可高效共享同一表征:在仅3.86%性能下降的情况下,自由观看训练模型可迁移至任务驱动的视觉搜索,以语义序列得分(SemSS)衡量预测注视路径与人类路径的相似性。该迁移使计算成本降低92.29%(GFLOPs)和31.23%(可训练参数)。

原文摘要 · Abstract (English)

Computational human attention modeling in free-viewing and task-specific settings is often studied separately, with limited exploration of whether a common representation exists between them. This work investigates this question and proposes a neural network architecture that builds upon the Human Attention transformer (HAT) to test the hypothesis. Our results demonstrate that free-viewing and visual search can efficiently share a common representation, allowing a model trained in free-viewing attention to transfer its knowledge to task-driven visual search with a performance drop of only 3.86% in the predicted fixation scanpaths, measured by the semantic sequence score (SemSS) metric which reflects the similarity between predicted and human scanpaths. This transfer reduces computational costs by 92.29% in terms of GFLOPs and 31.23% in terms of trainable parameters.

注意力模型知识迁移视觉搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。