用扩散模型生成多样真实的眼动轨迹,更好模拟人类视觉探索的变异性。
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
- 结合扩散模型与视觉变换器,利用随机性生成多样化眼动路径。
- 在基准数据集上优于现有方法,自由观看和任务驱动场景均提升多样性与准确性。
- 支持文本条件控制,可按任务目标生成对应眼动行为,适合人机交互研究。
预测人类眼动扫描路径对理解视觉注意至关重要,广泛应用于人机交互、自主系统和认知机器人领域。尽管深度学习模型已推动该方向发展,但多数方法仅生成平均行为,难以捕捉人类视觉探索的变异性。本文提出ScanDiff,一种融合扩散模型与视觉变换器的新架构,通过扩散模型的随机性显式建模眼动路径的多样性,生成大量合理的眼动轨迹。同时引入文本条件控制,实现任务驱动的眼动生成,使模型能适应不同视觉搜索目标。在多个基准数据集上的实验表明,ScanDiff在自由观看和任务驱动场景下均超越现有最先进方法,生成更丰富且准确的眼动路径。结果验证了其对人类视觉行为复杂性的更好建模能力,推动了眼动预测研究的发展。源代码与模型已公开发布于 https://aimagelab.github.io/ScanDiff。
原文摘要 · Abstract (English)
Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to capture the variability of human visual exploration. In this work, we present ScanDiff, a novel architecture that combines diffusion models with Vision Transformers to generate diverse and realistic scanpaths. Our method explicitly models scanpath variability by leveraging the stochastic nature of diffusion models, producing a wide range of plausible gaze trajectories. Additionally, we introduce textual conditioning to enable task-driven scanpath generation, allowing the model to adapt to different visual search objectives. Experiments on benchmark datasets show that ScanDiff surpasses state-of-the-art methods in both free-viewing and task-driven scenarios, producing more diverse and accurate scanpaths. These results highlight its ability to better capture the complexity of human visual behavior, pushing forward gaze prediction research. Source code and models are publicly available at https://aimagelab.github.io/ScanDiff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。