让第一人称视频理解具备3D空间感知能力,提升对物体位置关系的把握。
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
- 用伪深度图+轻量3D解码器学习视频的3D空间结构。
- 在多个下游任务中超越现有方法,尤其在空间关系理解上表现突出。
- 适合需要精准空间认知的智能助手、机器人视觉等应用。
第一人称视频-语言预训练显著推动了视频表征学习的发展。人类感知和交互的是一个完整的三维世界,具备超越文本理解的空间意识。然而,以往多数工作仅基于一维文本或二维视觉线索(如边界框)进行学习,缺乏真正的三维理解。为弥合这一差距,我们提出EgoDTM——一种结合大规模3D感知视频预训练与视频-文本对比学习的第一人称深度与文本感知模型。EgoDTM引入轻量级3D感知解码器,通过深度估计模型生成的伪深度图高效学习3D感知能力。为进一步支持3D感知视频预训练,我们通过有机融合多个基础模型,将原始简短描述丰富为包含手-物视觉线索的增强描述。大量实验表明,EgoDTM在多种下游任务中表现优异,凸显其更强的3D感知视觉理解能力。代码已开源:https://github.com/xuboshen/EgoDTM。
原文摘要 · Abstract (English)
Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most previous works learn from 1D text or 2D visual cues, such as bounding boxes, which inherently lack 3D understanding. To bridge this gap, we introduce EgoDTM, an Egocentric Depth- and Text-aware Model, jointly trained through large-scale 3D-aware video pretraining and video-text contrastive learning. EgoDTM incorporates a lightweight 3D-aware decoder to efficiently learn 3D-awareness from pseudo depth maps generated by depth estimation models. To further facilitate 3D-aware video pretraining, we enrich the original brief captions with hand-object visual cues by organically combining several foundation models. Extensive experiments demonstrate EgoDTM's superior performance across diverse downstream tasks, highlighting its superior 3D-aware visual understanding. Code: https://github.com/xuboshen/EgoDTM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。