arXiv:2412.04184cs.NEcs.LG2024-12被引 4

用频谱正则化GAN生成更真实的眨眼速度轨迹。

Modeling Eye Gaze Velocity Trajectories using GANs with Spectral Loss for Enhanced Fidelity

  • 采用LSTM-CNN结构结合频谱损失,捕捉眼动复杂时序特征。
  • 相比传统模型,生成数据在均值、标准差等统计量上更接近真实数据。
  • 适合人机交互、神经诊断等领域需要高保真眼动数据的研究者。

准确建模眼动动态对人机交互、神经诊断和认知研究至关重要。传统生成模型如隐马尔可夫模型(HMM)难以捕捉眼动轨迹中的复杂时序依赖和分布特性。本文提出一种基于GAN的框架,使用LSTM和CNN作为生成器与判别器,生成高保真合成眼动速度轨迹。评估了四种GAN架构:CNN-CNN、LSTM-CNN、CNN-LSTM、LSTM-LSTM,分别在仅使用对抗损失和加入加权对抗损失与频谱损失组合两种条件下训练。结果表明,使用频谱损失的LSTM-CNN架构最贴近真实数据分布,能有效捕捉分布尾部和精细时序依赖。频谱正则化显著提升模型对眼动频谱特性的还原能力,增强学习稳定性并提高数据保真度。与优化至4个隐藏状态的HMM对比显示,其生成数据在均值、标准差、偏度和峰度上显著偏离真实数据;而LSTM-CNN模型在这些统计量上与真实数据高度一致,验证了其对眼动动态复杂性的建模能力。该频谱正则化LSTM-CNN GAN可作为生成高质量合成眼动速度数据的可靠工具。

原文摘要 · Abstract (English)

Accurate modeling of eye gaze dynamics is essential for advancement in human-computer interaction, neurological diagnostics, and cognitive research. Traditional generative models like Markov models often fail to capture the complex temporal dependencies and distributional nuance inherent in eye gaze trajectories data. This study introduces a GAN framework employing LSTM and CNN generators and discriminators to generate high-fidelity synthetic eye gaze velocity trajectories. We conducted a comprehensive evaluation of four GAN architectures: CNN-CNN, LSTM-CNN, CNN-LSTM, and LSTM-LSTM trained under two conditions: using only adversarial loss and using a weighted combination of adversarial and spectral losses. Our findings reveal that the LSTM-CNN architecture trained with this new loss function exhibits the closest alignment to the real data distribution, effectively capturing both the distribution tails and the intricate temporal dependencies. The inclusion of spectral regularization significantly enhances the GANs ability to replicate the spectral characteristics of eye gaze movements, leading to a more stable learning process and improved data fidelity. Comparative analysis with an HMM optimized to four hidden states further highlights the advantages of the LSTM-CNN GAN. Statistical metrics show that the HMM-generated data significantly diverges from the real data in terms of mean, standard deviation, skewness, and kurtosis. In contrast, the LSTM-CNN model closely matches the real data across these statistics, affirming its capacity to model the complexity of eye gaze dynamics effectively. These results position the spectrally regularized LSTM-CNN GAN as a robust tool for generating synthetic eye gaze velocity data with high fidelity.

眼动建模GAN频谱损失生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。