arXiv:2608.08947cs.CVcs.HC2026-08

用摄像头注视追踪约束自动驾驶模型的虚假目标,发现精度不足无法实现。

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

  • 用驾驶者注视数据约束模型内部目标形成
  • 多组实验均未显著提升性能(p值均>0.5)
  • 摄像头注视误差远超目标物尺寸,无法定位

当前自动驾驶中的危险检测系统可能生成隐含目标,即通过虚假相关性而非真实危险识别来提升训练表现。我们探究是否可通过基于网络摄像头的眼动追踪(WebGazer.js)捕捉的人类注视模式,作为特权信息来约束隐含目标的形成。在388段真实行车记录仪视频中同步收集了137,663帧级注视样本与危险标注,测试了两种校准协议(9点/45点击和11点/440点击)、两种模型架构(随机森林与因果Transformer),每组实验重复5次并进行配对t检验。所有实验均未显示注视信息带来统计显著提升(p = 0.919, 0.578, 0.667)。几何分析表明根本原因:WebGazer报告误差(配置不同为130-257像素)超过93%已检测危险物体尺寸(中位数36像素),导致在该仪器精度下无法实现对象级注视归因。

原文摘要 · Abstract (English)

Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns, captured via webcam-based eye tracking (WebGazer.js), can serve as privileged information to constrain mesa-objective formation. We collected 137,663 frame-level gaze samples synchronized with hazard annotations across 388 real dashcam clips, then test this hypothesis across two calibration protocols (9-point/45-click and 11-point/440-click), two model architectures (Random Forest and causal Transformer), and five random seeds per experiment with paired t-tests. No experiment yields a statistically significant improvement from gaze (p = 0.919, 0.578, and 0.667 respectively). A geometric analysis reveals the root cause: WebGazer's reported error (~130-257 px depending on configuration) exceeds 93% of detected hazard object sizes (median 36 px), rendering object-level gaze attribution physically impossible at this instrument precision.

自动驾驶眼动追踪隐含目标模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。