arXiv:2603.28319cs.CV2026-03

用图模型模拟动态场景下人眼注意力,更真实还原注视轨迹。

Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic Scenes

  • 将注视建模为自回归动态系统,显式追踪时间序列轨迹。
  • 在30人数据集上生成更自然的注视路径与显著性图。
  • 适合研究驾驶安全、人机交互中的动态注意力建模。

准确建模人类注意力对计算机视觉应用至关重要,尤其在汽车安全领域。现有方法通常将注视简化为显著性图或注视路径,仅隐式处理动态。本文将注视建模为自回归动力系统,显式展开原始注视轨迹,基于注视历史与环境演化进行条件建模。驾驶场景以注视为中心的图结构表示,由异构图变压器Affinity Relation Transformer(ART)处理,建模驾驶员注视、交通物体与道路结构之间的交互。此外引入物体密度网络(ODN),预测下一步注视分布,捕捉复杂环境中注意力转移的随机性与对象中心特性。我们还发布了Focus100数据集,包含30名参与者观看第一视角驾驶视频的原始注视数据。直接在原始注视数据上训练,无需剔除固定点,所提统一方法生成的注视轨迹、注视路径动态和显著性图均优于现有模型,为动态环境下人类注意力的时序建模提供新洞见。

原文摘要 · Abstract (English)

Accurately modelling human attention is essential for numerous computer vision applications, particularly in the domain of automotive safety. Existing methods typically collapse gaze into saliency maps or scanpaths, treating gaze dynamics only implicitly. We instead formulate gaze modelling as an autoregressive dynamical system and explicitly unroll raw gaze trajectories over time, conditioned on both gaze history and the evolving environment. Driving scenes are represented as gaze-centric graphs processed by the Affinity Relation Transformer (ART), a heterogeneous graph transformer that models interactions between driver gaze, traffic objects, and road structure. We further introduce the Object Density Network (ODN) to predict next-step gaze distributions, capturing the stochastic and object-centric nature of attentional shifts in complex environments. We also release Focus100, a new dataset of raw gaze data from 30 participants viewing egocentric driving footage. Trained directly on raw gaze, without fixation filtering, our unified approach produces more natural gaze trajectories, scanpath dynamics, and saliency maps than existing attention models, offering valuable insights for the temporal modelling of human attention in dynamic environments.

注意力建模动态场景图神经网络驾驶安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。