用联合建模生成更自然的人类注视轨迹和注视路径。
ST-DiffEye: Diffusion-based Continuous Gaze Generation via Joint Scanpath-Trajectory Modeling

- 将注视轨迹与注视路径拼接输入,通过扩散模型联合生成。
- 在视觉搜索任务中达到当前最优性能,优于单一模态模型。
- 适合研究人眼行为、交互设计及生成式视觉模型的开发者。
本文研究人类注视建模问题,旨在生成观察视觉刺激时产生的注视模式。注视主要通过两种方式捕捉:连续眼动轨迹(描述精细运动动态)和离散注视路径(描述高层固定结构)。由于个体和实验间差异显著,我们将其视为本质特征而非噪声,将注视建模为随机生成过程。现有生成模型仅独立依赖其中一种表示进行监督。我们假设轨迹与路径在互补尺度上共同提供信息,提出ST-DiffEye——一种联合轨迹-路径扩散框架,通过拼接两者作为额外输入通道实现耦合,无需额外架构开销。此外,我们引入基于连续概率评分(CRPS)的严谨评估框架,将序列相似性指标推广为可衡量生成准确性与多样性的严格评分规则。在任务驱动的视觉搜索(含目标存在与不存在场景)及自由观看基准上的实验表明,该方法达到领先性能。详细消融分析证实联合建模的有效性以及分布感知评估对捕捉注视内在变异性的价值。
原文摘要 · Abstract (English)
We study the problem of human gaze modeling, which aims to generate the gaze patterns a viewer produces while observing a visual stimulus. Gaze is primarily captured through two modalities: continuous eye-tracking trajectories, which describe fine-grained motion dynamics, and discrete scanpaths, which describe high-level fixation structure. Because gaze varies substantially across viewers and trials, we treat this variability as a defining property rather than noise and model gaze as a stochastic generative process. Existing generative gaze models supervise on only one of these two representations in isolation. We hypothesize that trajectories and scanpaths describe gaze at complementary scales and are jointly informative during training, and test this hypothesis through ST-DiffEye, a joint trajectory-scanpath diffusion framework that couples both modalities by concatenating them as an additional raw input channel, requiring no architectural overhead beyond an input and output channel expansion. We further introduce a principled evaluation framework based on the Continuous Ranked Probability Score (CRPS), which generalizes any existing sequence similarity metric into a proper scoring rule that jointly assesses the accuracy and diversity of generated gaze. Experiments on task-driven visual search, covering both target-present and target-absent scenarios, and on free-viewing benchmarks demonstrate state-of-the-art performance. These results, along with detailed ablations, confirm the benefit of joint modeling and the value of distribution-aware evaluation in capturing the intrinsic variability of human gaze. Project webpage: https://st-diffeye.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。