用视线引导构建图结构,精准捕捉第一视角动作中的关键交互
G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding

- 以视线为线索构建动作场景图,筛选出与动作相关的目标实体
- 在EGTEA Gaze+和MECCANO数据集上达到媲美视频模型的准确率
- 无需昂贵视频预训练,适合处理类别不平衡的现实场景
第一人称动作理解通常依赖于在大量外部视角数据上预训练的大规模视频模型。然而,许多第一人称动作仅涉及少量手-物交互,且只与少数相关实体有关。本文提出G3Ego,一种基于图的框架,利用视线作为结构线索,识别场景中与动作相关的实体。从稀疏采样的帧中,G3Ego结合视觉-语言描述、可感知物体和手部线索构建动作场景图,并通过佩戴者视线剔除无关实体。最终的图嵌入经过时间聚合用于动作识别与预测。与以往将视线仅作为辅助模态或注意力信号不同,G3Ego将视线直接融入图结构构建,生成高效且可解释的表示,聚焦于动作相关交互。在EGTEA Gaze+和MECCANO数据集上的实验表明,G3Ego性能优于同类方法,在类别不平衡评估下持续提升Macro-F1,同时避免了计算开销巨大的视频预训练。结果证明视线引导图表示在第一人称动作理解中的有效性。
原文摘要 · Abstract (English)
Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving only a few relevant entities. We propose G3Ego, a graph-based framework for egocentric action understanding that uses gaze as a structural cue to identify action-relevant entities in the scene. From sparsely sampled frames, G3Ego constructs action scene graphs from vision-language descriptions, grounded objects, and hand cues, and then prunes irrelevant entities using the camera wearer's gaze. The resulting graph embeddings are temporally aggregated for action recognition and anticipation. Unlike prior work that uses gaze primarily as an auxiliary modality or attention signal, G3Ego incorporates gaze directly into graph construction, producing efficient and interpretable representations focused on action-relevant interactions. Experiments on EGTEA Gaze+ and MECCANO show that G3Ego achieves competitive performance compared with video-based approaches and consistently improves Macro-F1 under class-imbalanced evaluation, while avoiding reliance on computationally expensive video pretraining. These results demonstrate the effectiveness of gaze-guided graph representations for egocentric action understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。