arXiv:2603.24938cs.CV2026-03被引 1

用自回归扩散模型生成任意时长视频的连续眼动轨迹。

Infinite Gaze Generation for Videos with Autoregressive Diffusion

  • 基于视觉特征空间构建自回归扩散模型,生成连续坐标与高精度时间戳的眼动数据。
  • 在长时序预测上显著优于现有方法,能捕捉真实场景中长期行为依赖关系。
  • 适合关注视频理解、人机交互及长时间眼动建模的研究者。

预测视频中的人类注视行为是提升场景理解与多模态交互的关键。传统显著性图仅提供空间概率分布,而注视路径虽有序但常忽略原始注视的细粒度时序动态。现有模型通常局限于短时窗(约3-5秒),难以捕捉真实内容中的长程行为依赖。本文提出一种用于任意长度视频的无限时长原始注视预测生成框架。通过自回归扩散模型,我们合成具有连续空间坐标和高分辨率时间戳的注视轨迹。模型以显著性感知的视觉隐空间为条件。定量与定性评估表明,该方法在长时序时空精度与轨迹真实性方面显著优于现有方法。

原文摘要 · Abstract (English)

Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, both abstractions often collapse the fine-grained temporal dynamics of raw gaze. Furthermore, existing models are typically constrained to short-term windows ($\approx$ 3-5s), failing to capture the long-range behavioral dependencies inherent in real-world content. We present a generative framework for infinite-horizon raw gaze prediction in videos of arbitrary length. By leveraging an autoregressive diffusion model, we synthesize gaze trajectories characterized by continuous spatial coordinates and high-resolution timestamps. Our model is conditioned on a saliency-aware visual latent space. Quantitative and qualitative evaluations demonstrate that our approach significantly outperforms existing approaches in long-range spatio-temporal accuracy and trajectory realism.

眼动预测扩散模型视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。