提出分层框架,让系统同时理解视频中多时间尺度的行为
Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives
- 设计分层结构与专用图神经网络,处理不同时间粒度的推理
- 在Ego4d多个基准上实现剪辑级与帧级任务的联合准确率提升
- 适合需要多粒度行为理解的智能视觉系统研发人员
人类对动作视频的理解是多维度的:几秒内即可把握当前事件、识别物体间关联与交互,并预测后续发展。为使自主系统具备这种整体感知能力,需学会跨任务关联概念、抽象知识并利用任务协同学习新技能。此前的EgoPack框架已实现多种任务间的高效信息共享。本文提出Hier-EgoPack,通过引入专为多时间粒度设计的图神经网络层,扩展了该框架的时空推理能力。我们在包含剪辑级和帧级推理的多个Ego4d基准上验证方法,结果表明该分层统一架构可同时有效解决多样化任务。
原文摘要 · Abstract (English)
Our comprehension of video streams depicting human activities is naturally multifaceted: in just a few moments, we can grasp what is happening, identify the relevance and interactions of objects in the scene, and forecast what will happen soon, everything all at once. To endow autonomous systems with such a holistic perception, learning how to correlate concepts, abstract knowledge across diverse tasks, and leverage tasks synergies when learning novel skills is essential. A significant step in this direction is EgoPack, a unified framework for understanding human activities across diverse tasks with minimal overhead. EgoPack promotes information sharing and collaboration among downstream tasks, essential for efficiently learning new skills. In this paper, we introduce Hier-EgoPack, which advances EgoPack by enabling reasoning also across diverse temporal granularities, which expands its applicability to a broader range of downstream tasks. To achieve this, we propose a novel hierarchical architecture for temporal reasoning equipped with a GNN layer specifically designed to tackle the challenges of multi-granularity reasoning effectively. We evaluate our approach on multiple Ego4d benchmarks involving both clip-level and frame-level reasoning, demonstrating how our hierarchical unified architecture effectively solves these diverse tasks simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。