arXiv:2501.04121cs.CV2025-01

用图模型对齐第一人称与第三人称视频,提升动作关键步骤识别准确率。

Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition

  • 构建视频片段为节点的图结构,利用长时依赖关系建模。
  • 在Ego-Exo4D上提升超过12个百分点,达到90.3%准确率。
  • 支持多模态信息融合,适合复杂动作识别任务研究者。

第一人称视频因动态背景、频繁运动和遮挡,给关键步骤识别带来挑战。本文提出一种灵活的图学习框架,通过构建以每个视频片段为节点的图结构,有效建模第一人称视频中的长时依赖,并在训练中利用第一人称与第三人称视频间的对齐关系,提升第一人称视频的推理性能。训练时将每个第三人称视频片段也作为额外节点加入图中。探索多种节点连接策略,将关键步骤识别转化为图上的节点分类任务。在Ego-Exo4D数据集上进行大量实验,结果表明该方法显著优于现有方法,准确率提升超过12个百分点。所构建图结构稀疏且计算高效。此外,还研究了语音、深度图和物体类别标签等多模态特征在异构图中的作用,分析其对识别性能的贡献。

原文摘要 · Abstract (English)

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning framework for fine-grained keystep recognition that is able to effectively leverage long-term dependencies in egocentric videos, and leverage alignment between egocentric and exocentric videos during training for improved inference on egocentric videos. Our approach consists of constructing a graph where each video clip of the egocentric video corresponds to a node. During training, we consider each clip of each exocentric video (if available) as additional nodes. We examine several strategies to define connections across these nodes and pose keystep recognition as a node classification task on the constructed graphs. We perform extensive experiments on the Ego-Exo4D dataset and show that our proposed flexible graph-based framework notably outperforms existing methods by more than 12 points in accuracy. Furthermore, the constructed graphs are sparse and compute efficient. We also present a study examining on harnessing several multimodal features, including narrations, depth, and object class labels, on a heterogeneous graph and discuss their corresponding contribution to the keystep recognition performance.

动作识别图神经网络第一人称视频多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。