通过挖掘人类行为层次结构,提升第一视角视频的推理能力
HiERO: understanding the hierarchy of human behavior enhances reasoning on egocentric videos
- 从无剧本视频中自动学习行为层级结构,增强视频特征
- 在多个基准上实现领先性能,零样本下比全监督方法高12.5% F1
- 适合做第一视角视频理解与步骤推理的研究者和开发者
人类活动高度复杂且多变,使深度学习模型难以进行有效推理。但我们发现这种多样性背后存在由相关动作组成的层级结构。本文提出HiERO,一种弱监督方法,通过将视频片段与叙述描述对齐,利用层级架构推断上下文、语义和时间关系,从而丰富视频片段的特征表示。在EgoMCQ、EgoNLQ等多个视频文本对齐基准上,仅需极少额外训练即取得优异表现;在零样本程序学习任务(EgoProceL和Ego4D Goal-Step)中同样超越现有方法,尤其在EgoProceL上零样本表现比全监督方法高出12.5% F1。结果表明,利用人类行为的层次知识可显著提升第一视角视觉中的多种推理任务性能。
原文摘要 · Abstract (English)
Human activities are particularly complex and variable, and this makes challenging for deep learning models to reason about them. However, we note that such variability does have an underlying structure, composed of a hierarchy of patterns of related actions. We argue that such structure can emerge naturally from unscripted videos of human activities, and can be leveraged to better reason about their content. We present HiERO, a weakly-supervised method to enrich video segments features with the corresponding hierarchical activity threads. By aligning video clips with their narrated descriptions, HiERO infers contextual, semantic and temporal reasoning with an hierarchical architecture. We prove the potential of our enriched features with multiple video-text alignment benchmarks (EgoMCQ, EgoNLQ) with minimal additional training, and in zero-shot for procedure learning tasks (EgoProceL and Ego4D Goal-Step). Notably, HiERO achieves state-of-the-art performance in all the benchmarks, and for procedure learning tasks it outperforms fully-supervised methods by a large margin (+12.5% F1 on EgoProceL) in zero shot. Our results prove the relevance of using knowledge of the hierarchy of human activities for multiple reasoning tasks in egocentric vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。