arXiv:2604.05475cs.CV2026-04

用3D模拟器生成可自动标注的眼动视频,用于检测视频面试中的读稿行为。

A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator

论文配图:A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator
图 1 · 摘自论文原文
  • 从真实视频提取人眼轨迹,通过无头浏览器在3D模拟器中重放生成眼动数据。
  • 构建了144段共12小时的合成眼动视频,包含72段读稿和72段对话场景。
  • 生成数据保留原始轨迹的时间动态,适合视觉语言模型与行为分析研究。

大规模视觉语言模型虽在互联网数据上表现优异,但行为模态(如眼动、手势)数据仍稀缺且难标注。为此,我们提出一种基于3D眼动模拟器的合成眼动视频生成方案:从参考视频提取真实人眼虹膜轨迹,并通过无头浏览器自动化技术在3D环境中重放。应用于视频面试中的脚本读取检测任务,我们发布了final_dataset_v1:共144个会话(72个读稿,72个对话),总计12小时合成眼动视频,帧率25fps。评估显示,生成轨迹在时间动态上与源数据高度一致(所有指标下KS D < 0.14)。帧级对比表明,该模拟器在读稿尺度运动上存在有限敏感性,主因是未建模头部协同运动,此发现有助于改进未来模拟器设计。整个管道、数据集及评估工具已开源,支持行为分类器在视觉语言系统中的开发。

原文摘要 · Abstract (English)

Large vision-language models have achieved remarkable capabilities by training on massive internet-scale data, yet a fundamental asymmetry persists: while LLMs can leverage self-supervised pretraining on abundant text and image data, the same is not true for many behavioral modalities. Video-based behavioral data -- gestures, eye movements, social signals -- remains scarce, expensive to annotate, and privacy-sensitive. A promising alternative is simulation: replace real data collection with controlled synthetic generation to produce automatically labeled data at scale. We introduce infrastructure for this paradigm applied to eye movement, a behavioral signal with applications across vision-language modeling, virtual reality, robotics, accessibility systems, and cognitive science. We present a pipeline for generating synthetic labeled eye movement video by extracting real human iris trajectories from reference videos and replaying them on a 3D eye movement simulator via headless browser automation. Applying this to the task of script-reading detection during video interviews, we release final_dataset_v1: 144 sessions (72 reading, 72 conversation) totaling 12 hours of synthetic eye movement video at 25fps. Evaluation shows that generated trajectories preserve the temporal dynamics of the source data (KS D < 0.14 across all metrics). A matched frame-by-frame comparison reveals that the 3D simulator exhibits bounded sensitivity at reading-scale movements, attributable to the absence of coupled head movement -- a finding that informs future simulator design. The pipeline, dataset, and evaluation tools are released to support downstream behavioral classifier development at the intersection of behavioral modeling and vision-language systems.

眼动仿真行为建模合成数据视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。