用运动特征构建眼动数据的紧凑表示,提升跨数据集迁移效果
Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

- 以事件对齐的角位移定义可解释的运动词汇表
- 在特定迁移方向上,运动令牌比固定码本表现更优
- 适合关注眼动跨数据集建模的研究者
眼动信号日益作为视觉与多模态模型的输入,但其跨数据集表示尚无共识。原始轨迹保留细节但噪声大且依赖设备,粗粒度事件标签易建模却可能丢失局部运动结构。本文提出事件对齐、固定时窗的角位移作为可解释的事件条件运动词汇,并与仅事件、空间、绝对角度、学习型向量量化及连续表示对比。评估结合下一令牌预测与目标域后悔值、低阶目标参考、配对自举、顺序敏感性、模式重叠和冻结结构探针。在事件对齐头戴设备基准测试中,角运动令牌在某一迁移方向上的目标域后悔值低于冻结码本VQ令牌,反向则不显著。探针揭示互补表示特性,仅事件令牌显示低困惑度仍可能丢失运动信息。在第三个远端数据集上,匹配比较I-VT、原生事件与帧跨度接口发现,事件构建方式显著影响迁移:原生事件在迁入EGTEA时后悔值最低,而帧跨度事件零模式重叠,在源任务上严重失效。因此,基于运动的标记化为事件对齐的自指眼动流提供了紧凑表示,评估亦揭示目标可预测性与事件构建如何塑造跨数据集结论。
原文摘要 · Abstract (English)
Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse event labels are easy to model but can discard local motion structure. We formulate event-aligned, fixed-horizon angular displacement as an interpretable, event-conditioned motion vocabulary and compare it with event-only, spatial, absolute-angle, learned vector-quantized, and continuous representations. To assess transfer alongside target predictability and token collapse, our evaluation combines next-token prediction with target-domain regret, low-order target references, paired bootstrap, order sensitivity, motif overlap, and frozen structural probes. In an event-aligned headset benchmark, angular-motion tokens have lower target-domain regret than frozen-codebook VQ tokens in one transfer direction, while the reverse direction is inconclusive. The probes reveal complementary representation properties, and event-only tokens show that low perplexity can retain little motion information. On a third egocentric dataset, a matched comparison of I-VT, native, and frame-span interfaces shows that event construction materially changes transfer: native events have the lowest regret into EGTEA, while frame-span events have zero motif overlap and fail severely as a source. Motion-based tokenization therefore provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。