arXiv:2509.16767cs.CV2025-09NeurIPS被引 5

用扩散模型生成自然图像下的连续眼动轨迹,更真实还原人类视觉注意力多样性。

DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images

  • 基于扩散模型建模连续眼动轨迹,引入位置对齐嵌入增强视觉特征匹配
  • 在小规模数据上生成高质量眼动模式,能还原真实注视分布的多样性
  • 首次实现自然图像上连续眼动轨迹生成,适合视觉认知与人机交互研究

大量眼动预测模型通常基于离散注视点序列(scanpath)训练,忽略了原始眼动轨迹中的丰富信息。现有方法多生成固定长度、单一路径的扫描路径,难以体现人类观看同一图像时的个体差异和随机性。为此,我们提出 DiffEye,一种基于扩散模型的连续眼动轨迹生成框架,以自然图像为条件。该方法引入对应位置嵌入(CPE),将空间注视信息与图像块语义特征对齐。通过使用原始眼动轨迹而非扫描路径进行训练,DiffEye 能捕捉人类注视行为的内在多样性,在较小数据集上仍生成高保真眼动模式。生成的轨迹可转换为扫描路径与显著性图,输出更符合真实人类视觉注意力分布。DiffEye 是首个在自然图像上利用扩散模型生成连续眼动轨迹的方法,实验表明其在扫描路径生成上达到当前最佳性能,并首次实现连续眼动轨迹生成。

原文摘要 · Abstract (English)

Numerous models have been developed for scanpath and saliency prediction, which are typically trained on scanpaths, which model eye movement as a sequence of discrete fixation points connected by saccades, while the rich information contained in the raw trajectories is often discarded. Moreover, most existing approaches fail to capture the variability observed among human subjects viewing the same image. They generally predict a single scanpath of fixed, pre-defined length, which conflicts with the inherent diversity and stochastic nature of real-world visual attention. To address these challenges, we propose DiffEye, a diffusion-based training framework designed to model continuous and diverse eye movement trajectories during free viewing of natural images. Our method builds on a diffusion model conditioned on visual stimuli and introduces a novel component, namely Corresponding Positional Embedding (CPE), which aligns spatial gaze information with the patch-based semantic features of the visual input. By leveraging raw eye-tracking trajectories rather than relying on scanpaths, DiffEye captures the inherent variability in human gaze behavior and generates high-quality, realistic eye movement patterns, despite being trained on a comparatively small dataset. The generated trajectories can also be converted into scanpaths and saliency maps, resulting in outputs that more accurately reflect the distribution of human visual attention. DiffEye is the first method to tackle this task on natural images using a diffusion model while fully leveraging the richness of raw eye-tracking data. Our extensive evaluation shows that DiffEye not only achieves state-of-the-art performance in scanpath generation but also enables, for the first time, the generation of continuous eye movement trajectories. Project webpage: https://diff-eye.github.io/

眼动追踪扩散模型连续轨迹注意力生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。