建模手物细粒度动态,提升第一人称视频理解效果
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
- 用检测器+大模型生成带手物交互描述的高质量数据
- 提出轻量级运动适配器,捕捉精细手物运动信息
- 在零样本任务上提升超16%,适合机器人交互研究
在第一人称视频理解中,手与物体的运动及其交互至关重要。然而现有方法主要关注视频表征与高层叙述对齐,忽视了手物之间的复杂动态。本文提出将细粒度手物动态建模融入视频表征学习流程。由于缺乏合适数据,我们构建HOD数据生成管道,结合手物检测器与大语言模型,生成包含详细手物动态描述的高质量叙述。为学习这些动态,提出EgoVideo模型,引入轻量级运动适配器以捕捉精细的手物运动信息。通过联合训练策略,EgoVideo有效利用了HOD数据中的细粒度动态。大量实验表明,该方法在多个第一人称下游任务中达到领先性能:在EK-100多实例检索上提升6.3%,分类任务提升5.7%,在EGTEA零样本分类上提升16.3%。此外,模型在手物交互与机器人操作任务中展现出强泛化能力。代码与数据已开源。
原文摘要 · Abstract (English)
In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate dynamics between hands and objects. In this work, we aim to integrate the modeling of fine-grained hand-object dynamics into the video representation learning process. Since no suitable data is available, we introduce HOD, a novel pipeline employing a hand-object detector and a large language model to generate high-quality narrations with detailed descriptions of hand-object dynamics. To learn these fine-grained dynamics, we propose EgoVideo, a model with a new lightweight motion adapter to capture fine-grained hand-object motion information. Through our co-training strategy, EgoVideo effectively and efficiently leverages the fine-grained hand-object dynamics in the HOD data. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple egocentric downstream tasks, including improvements of 6.3% in EK-100 multi-instance retrieval, 5.7% in EK-100 classification, and 16.3% in EGTEA classification in zero-shot settings. Furthermore, our model exhibits robust generalization capabilities in hand-object interaction and robot manipulation tasks. Code and data are available at https://github.com/OpenRobotLab/EgoHOD/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。