arXiv:2602.14837cs.CV2026-02TPAMI

融合环境可用性与注意力机制,提升短时物体交互预测准确率。

Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation

  • 设计双注意力架构,结合多尺度特征与时序池化,从图像对中预测交互
  • 在Ego4D上提升23个百分点,在EPIC-Kitchens上提升31个百分点
  • 适合可穿戴设备和人机协作场景,代码数据已开源

短时物体交互预测旨在通过观察第一人称视频,识别下一个活跃物体的位置、交互的名词与动词类别,以及接触时间。该能力对可穿戴助手理解用户目标并及时提供帮助,或实现人机交互至关重要。本文提出STAformer与STAformer++两种新型基于注意力的架构,融合帧引导时序池化、双图像-视频注意力及多尺度特征融合,支持从图像输入的视频对中进行预测。此外,引入两个新模块,通过建模可用性来锚定预测结果:一是环境可用性模型,作为特定物理场景中可能交互的持续记忆;二是基于手部与物体轨迹预测交互热点,增强热点区域预测置信度。实验表明,在Ego4D上总体Top-5 mAP提升最高达+23p.p,EPIC-Kitchens新标注集上提升+31p.p。代码、标注及预提取的可用性数据已在Ego4D和EPIC-Kitchens公开,以促进该领域研究。

原文摘要 · Abstract (English)

Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to contact from the observation of egocentric video. This ability is fundamental for wearable assistants to understand user goals and provide timely assistance, or to enable human-robot interaction. In this work, we present a method to improve the performance of STA predictions. Our contributions are two-fold: 1 We propose STAformer and STAformer plus plus, two novel attention-based architectures integrating frame-guided temporal pooling, dual image-video attention, and multiscale feature fusion to support STA predictions from an image-input video pair; 2 We introduce two novel modules to ground STA predictions on human behavior by modeling affordances. First, we integrate an environment affordance model which acts as a persistent memory of interactions that can take place in a given physical scene. We explore how to integrate environment affordances via simple late fusion and with an approach which adaptively learns how to best fuse affordances with end-to-end predictions. Second, we predict interaction hotspots from the observation of hands and object trajectories, increasing confidence in STA predictions localized around the hotspot. Our results show significant improvements on Overall Top-5 mAP, with gain up to +23p.p on Ego4D and +31p.p on a novel set of curated EPIC-Kitchens STA labels. We released the code, annotations, and pre-extracted affordances on Ego4D and EPIC-Kitchens to encourage future research in this area.

交互预测注意力机制可用性建模可穿戴系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。