通过虚拟现实实验,研究了行人眼动对自动驾驶接驳车场景下轨迹预测的提升作用。
Eye Gaze-Informed and Context-Aware Pedestrian Trajectory Prediction in Shared Spaces with Automated Shuttles: A Virtual Reality Study
- 融合眼动、头朝向与情境信息的多模态预测模型,分析眼动的贡献机制。
- 在锐角接近时眼动比头朝向提供额外预测信息,可降低8.47%的轨迹误差。
- 眼动与情境信息互补,适合用于人机共驾环境中的行为预测系统设计。
为填补这一空白,我们开展了一项虚拟现实实验,让行人与自动驾驶接驳车在不同接近角度(45°、90°、135°)及连续交通条件(单辆接驳车、两辆接驳车间隔3秒或5秒)下交互,同步采集运动、眼动和头朝向数据。为探究细粒度眼动在何种条件下、以何种形式有助于行人轨迹预测,我们构建了一个多模态预测模型,通过模态专用编码器融合这些信号,并系统性消融眼动表征与头朝向、情境上下文的关系。主要发现:第一,眼动的预测价值具有角度依赖性,且与眼-头-躯干协同密切相关;在锐角接近时,行人主动调整视线以获取接驳车,此时眼动包含头朝向无法捕捉的信息。第二,持续的眼动方向优于分类化的语义注视标签,最优编码方式(全局或身体相对)取决于眼动是否单独使用或与情境联合使用。第三,眼动与情境信息提供互补性预测信息,两者结合可使最终位移误差(FDE)降低8.47%,接近各自独立贡献之和。这些结果凸显将人类感知信号融入行人行为预测的重要性,推动发展以人为核心的补充性建模方法。代码已开源:https://github.com/danyayay/GazeX.git。
原文摘要 · Abstract (English)
To address this gap, we conduct a Virtual Reality experiment in which pedestrians interact with automated shuttles under varying approach angles (45°, 90°, 135°) and continuous-traffic conditions (single shuttle, two shuttles with 3 or 5-second gaps), collecting synchronized motion, eye gaze, and head orientation data. To investigate to what extent, under what conditions, and in what form fine-grained eye gaze is informative for pedestrian motion prediction, we develop a multi-modal prediction model that fuses these signals through modality-specific encoders, and systematically ablate gaze representations against head orientation and situational context. We report three main results. First, the predictive value of eye gaze is angle-dependent and tightly coupled with eye-head-body coordination: at acute angles where pedestrians actively redirect gaze to acquire the shuttle, eye gaze carries information that head orientation alone misses. Second, continuous gaze orientation outperforms categorical semantic fixation labels, with the optimal encoding frame (global or body-relative) depending on whether gaze is used alone or jointly with context. Third, eye gaze and situational context provide complementary predictive information: their combination reduces final displacement error (FDE) by 8.47%, close to the sum of their individual contributions. Together, these findings highlight the value of incorporating human perceptual signals into pedestrian behavior prediction and motivate a human-centered complement to vehicle-centric modeling approaches. Our code is available at https://github.com/danyayay/GazeX.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。