研究未来信息对视线预测的帮助,发现适量未来数据能显著提升效果。
How Much Future Helps? A Controlled Study of Future-Privileged Supervision for Causal Egocentric Gaze Estimation

- 用可调未来窗口训练模型,推理时保持因果性
- 最佳效果出现在1.7~3.3秒未来上下文(EGTEA Gaze+)和2.7秒(Ego4D)
- 为实时视线建模提供实用训练指导
视角自注视估计通常在模型可访问完整视频(含未来帧)的条件下进行研究,而真实应用需严格因果、在线预测。这引发关键问题:未来上下文是否本质上提供有价值信号?若然,多长的未来预览在训练中最优?为此,我们提出一个受控框架,包含一个训练时可访问可调未来视窗的分支,推理时被移除。该设计隔离了未来上下文的影响,同时保持推理架构严格因果。在EGTEA Gaze+和Ego4D数据集上,我们发现未来特权监督始终提升因果预测性能,证实其有效性。然而,性能增益不随未来窗口延长单调上升,而是在有限时间范围内达到峰值:在EGTEA Gaze+上为1.7–3.3秒($H{ imes}[5, 10]$),在Ego4D上为2.7秒($H{=}10$)。结果表明,轻量级因果模型可有效吸收未来感知信号,为实时视角自注视建模提供实践指导。
原文摘要 · Abstract (English)
Egocentric gaze estimation is commonly studied using models that process the full video with access to future frames, while real-world applications require strictly causal, online prediction. This discrepancy raises key questions: Does future context inherently provide valuable signals for gaze estimation? If so, how much future look-ahead optimally supervises a causal model during training? To investigate, we propose a controlled framework featuring a future-aware branch that accesses a tunable look-ahead horizon during training but is discarded at inference. This design isolates the impact of future context while keeping the inference architecture fixed and strictly causal. Across EGTEA Gaze+ and Ego4D, we find that future-privileged supervision consistently improves causal gaze prediction, confirming its utility. However, performance gains do not increase monotonically with longer look-ahead, but rather peak within a bounded temporal regime. Specifically, optimal performance corresponds to roughly 1.7--3.3 seconds of future context ($H{\in}[5, 10]$) on EGTEA Gaze+ and 2.7 seconds ($H{=}10$) on Ego4D. Our results demonstrate that lightweight causal models can effectively absorb future-aware signals, providing practical guidance for real-time egocentric gaze modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。