arXiv:2609.05522cs.CVcs.AI2026-09

用扩散模型生成眼动轨迹,省去昂贵数据采集。

Diffusion models for eye-gaze trajectory generation using position and velocity representations

论文配图:Diffusion models for eye-gaze trajectory generation using position and velocity representations
图 1 · 摘自论文原文
  • 设计双扩散模型,分别生成位置与速度轨迹序列。
  • 位置模型JS散度仅0.016±0.004,固定时长误差<2%。
  • 适合需合成眼动数据的研究者,如人机交互、心理学实验。

眼动数据采集成本高且受隐私限制难以共享。本文采用两种互补的去噪扩散概率模型(DDPM),基于视觉搜索数据无条件生成眼动动态。两者均使用相同的FiLM条件一维U-Net(1935万参数),在28名参与者的数据中训练8秒滑动窗口序列。一个模型生成二维眼动位置序列,另一个生成两分量速度序列;各自采用特定预处理、训练设置、数据划分和评估协议。在三个独立训练种子下评估,报告均值±标准差。位置空间模型在九个运动学特征上平均JS散度为0.016±0.004,最高特征均值低于0.030,固定时长误差小于2%,弗雷歇眼动距离比统计与马尔可夫基线低一个数量级。在仅合成数据训练、真实数据测试的协议下,达到R²=0.66±0.02,为真实数据基准的82.7%。速度空间模型在速度、对数速度和转向角上平均JS散度为0.0065,最大值0.015±0.005。重构路径长度准确性较差(0.21±0.02对比位置空间0.03±0.01),但评估协议不同。总体表明,无条件扩散模型能捕捉局部眼动动力学及短程时空结构,而长程特性如扫视次数与累积路径几何仍需未来条件化模型解决。

原文摘要 · Abstract (English)

Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints. We address this using two complementary denoising diffusion probabilistic models (DDPMs) for unconditional generation of eye-gaze dynamics from visual-search data. Both use an identical FiLM-conditioned one-dimensional U-Net with self-attention (19.35,M parameters), trained on 8,s sliding-window sequences from 28 participants. One model generates raw two-dimensional gaze-position sequences, while the other generates two-component velocity sequences; each uses representation-specific preprocessing, training settings, data partitions, and evaluation protocols. Both are evaluated across three independent training seeds, with aggregated metrics reported as mean,$\pm$,SD. The position-space model achieves a mean Jensen-Shannon (JS) divergence of $0.016\pm0.004$ across nine kinematic features, with the highest feature-wise mean below $0.030$, fixation duration within 2% of real data, and a Fr'echet Gaze Distance more than an order of magnitude below statistical and Markovian baselines. Under a Train-on-Synthetic-Test-on-Real protocol, synthetic-only training achieves $R^2=0.66\pm0.02$, or 82.7% of the real-data $R^2$ point estimate. The velocity-space model achieves a mean JS divergence of $0.0065$ across velocity components, speed, log-speed, and turning angle, with a maximum of $0.015\pm0.005$. Reconstructed path length is less accurate ($0.21\pm0.02$ versus $0.03\pm0.01$ in position space), although the protocols differ. Overall, unconditional diffusion captures local gaze kinematics and short-range temporal and directional structure, while long-range properties such as saccade counts and cumulative path geometry remain targets for future conditioned models.

眼动生成扩散模型轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。