arXiv:2601.13218cs.CV2026-01中稿 · publication at the…

构建首个面向交互式过街场景的物体级视觉注意力数据集,推动注意力模型精准预测。

ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments

  • 基于虚拟现实采集120人过街行为,标注物体级注视数据与环境状态。
  • 提出oSIM新指标,显式优化物体注意力可提升模型在主流指标上的表现。
  • 设计SUMGraph模型,用图结构显式建模关键物体,性能超越现有方法。

人类视觉注意力的物体基础特性在认知科学中已有广泛认知,但在计算视觉注意力模型中应用有限,主要受限于缺乏合适的数据集和评估指标。为此,本文提出ObjectVisA-120——一个包含120名参与者在虚拟现实环境中进行空间过街导航的数据集,专为物体级注意力评估设计。该数据集因涉及伦理与安全挑战,难以在真实环境中获取。ObjectVisA-120不仅包含精确的眼动追踪数据及虚拟环境中物体的完整状态空间表示,还提供可变场景复杂度与丰富标注,包括全景分割、深度信息和车辆关键点。我们进一步提出物体级相似性(oSIM)作为新型评估指标,用于衡量物体级注意力模型性能,此前未被探索。实验表明,显式优化物体注意力不仅能提升oSIM表现,还能改善模型在传统指标上的效果。此外,本文提出SUMGraph,一种基于Mamba U-Net的模型,通过图结构显式编码关键场景物体(如车辆),在多个最先进的视觉注意力预测方法上实现性能提升。数据集、代码与模型将公开发布。

原文摘要 · Abstract (English)

The object-based nature of human visual attention is well-known in cognitive science, but has only played a minor role in computational visual attention models so far. This is mainly due to a lack of suitable datasets and evaluation metrics for object-based attention. To address these limitations, we present ObjectVisA-120 -- a novel 120-participant dataset of spatial street-crossing navigation in virtual reality specifically geared to object-based attention evaluations. The uniqueness of the presented dataset lies in the ethical and safety affiliated challenges that make collecting comparable data in real-world environments highly difficult. ObjectVisA-120 not only features accurate gaze data and a complete state-space representation of objects in the virtual environment, but it also offers variable scenario complexities and rich annotations, including panoptic segmentation, depth information, and vehicle keypoints. We further propose object-based similarity (oSIM) as a novel metric to evaluate the performance of object-based visual attention models, a previously unexplored performance characteristic. Our evaluations show that explicitly optimising for object-based attention not only improves oSIM performance but also leads to an improved model performance on common metrics. In addition, we present SUMGraph, a Mamba U-Net-based model, which explicitly encodes critical scene objects (vehicles) in a graph representation, leading to further performance improvements over several state-of-the-art visual attention prediction methods. The dataset, code and models will be publicly released.

视觉注意力虚拟现实图神经网络数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。