arXiv:2607.04017cs.CV2026-07中稿 · ECCV

统一建模动作与视线,实现行为理解的实时识别与预测。

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

论文配图:SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding
图 1 · 摘自论文原文
  • 用Transformer架构融合时空注意力,同步预测当前和未来动作与视线。
  • 在三个数据集上表现超越单一任务专用模型,最高提升12.3%准确率。
  • 适用于第一视角和第三人称场景,推动人机交互更自然智能。

人类与物体的交互(HOI)、视线模式及其预测紧密关联,为认知过程、意图和行为理解提供关键线索。然而,现有模型多将视线与动作分开处理,忽略了二者相互依赖性及统一建模的优势。本文提出新型统一框架SAGE(Synchronized Action-GazE),将HOI与人类视线的同步识别与预测集成于单一端到端可训练模型中。该方法基于Transformer架构,将视线信息融入时空注意力机制,同时预测当前与未来的动作和视线行为。我们在不同场景下探索了视线与动作间的双向关系,涵盖近距离细节视图(第一人称)和远距离上下文视图(第三人称),使框架具备广泛应用灵活性。此外,由于缺乏支持第三人称视频中全面分析HOI与视线的数据集,我们构建了新基准Exo-Cook以促进该领域研究。在三个基准数据集——VidHOI、EGTEA Gaze+和Exo-Cook上的实验表明,联合建模当前与未来帧中的动作与视线,能持续取得优异性能,常优于针对单任务优化的顶尖模型。通过综合统一动作与注意力,本工作为更直观的人机交互奠定基础。

原文摘要 · Abstract (English)

Human object interaction (HOI), gaze pattern, and their anticipation are intricately linked, providing valuable insights into cognitive processes, intentions, and behavior. However, most existing models handle gaze and actions separately, missing both their interdependence and the advantages of a unified solution. This paper presents a novel unified framework, SAGE (Synchronized Action-GazE), which integrates simultaneous recognition and anticipation of both HOI and human gaze into a single unified end-to-end trainable model. Our approach leverages a transformer-based architecture and incorporates gaze data into spatiotemporal attention mechanisms to simultaneously predict current and future human actions and gaze behavior. We explore this bidirectional relationship between gaze and actions under different scenarios, whether requiring a close-up, detailed view (egocentric) or a wider, more contextual view (exocentric), making our framework versatile for various applications. Additionally, due to lack of datasets for comprehensive analysis of both HOI and gaze in exocentric videos, we establish a new benchmark Exo-Cook to facilitate further research in this domain. Experiments on three benchmark datasets: VidHOI, EGTEA Gaze+, and Exo-Cook show that jointly modeling gaze and actions across current and future frames achieves consistently strong results, often surpassing specialized state-of-the-art models tailored to individual tasks. By unifying actions and attention in a comprehensive way, our work lays the groundwork for more intuitive human-machine interaction.

行为理解视线预测统一建模Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。