arXiv:2506.11558cs.CVcs.AI2025-06

DaMO模型提升视频语言模型的精准时间推理能力,尤其在弱监督下表现优异。

DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs

  • 采用分层双流结构融合视觉与音频信息,增强时间动态捕捉
  • 在多个基准上超越现有方法,尤其在精确时间定位任务中优势显著
  • 适合需要高精度时间推理的应用,如智能视频分析与教育场景

大型语言模型(LLMs)已拓展至视频领域,实现复杂的视频-语言理解。然而,现有视频LLM在细粒度时间推理方面存在局限,难以精确将回答关联到具体视频时刻,尤其是在监督受限条件下。本文提出DaMO,一种专为准确时间推理和多模态理解设计的数据高效视频LLM。其核心是时间感知的Fuseformer,采用分层双流架构,逐步捕捉各模态内的时序动态,并有效融合互补的视觉与音频信息。为提升计算效率,DaMO引入全局残差机制,减少空间冗余同时保留关键语义细节。通过四阶段渐进式训练范式,逐步赋予模型多模态对齐、语义定位和时间推理能力。本工作还基于现有数据集,利用大模型生成带时间锚定的问答对,构建多个增强数据集,用于需要时间标注的任务。在时间定位和视频问答基准上的全面实验表明,DaMO持续优于先前方法,尤其在要求精确时间对齐与推理的任务中表现突出。该研究为数据高效的视频-语言建模提供了新方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently been extended to the video domain, enabling sophisticated video-language understanding. However, existing Video LLMs often exhibit limitations in fine-grained temporal reasoning, restricting their ability to precisely attribute responses to specific video moments, especially under constrained supervision. We introduce DaMO, a data-efficient Video LLM explicitly designed for accurate temporal reasoning and multimodal understanding. At its core, the proposed Temporal-aware Fuseformer employs a hierarchical dual-stream architecture that progressively captures temporal dynamics within each modality and effectively fuses complementary visual and audio information. To further enhance computational efficiency, DaMO integrates a global residual that reduces spatial redundancy while preserving essential semantic details. We train DaMO via a structured four-stage progressive training paradigm, incrementally equipping the model with multimodal alignment, semantic grounding, and temporal reasoning capabilities. This work also contributes multiple datasets augmented from existing ones with LLM-generated temporally grounded QA pairs for tasks requiring temporal supervision. Comprehensive experiments on temporal grounding and video QA benchmarks demonstrate that DaMO consistently surpasses prior methods, particularly in tasks demanding precise temporal alignment and reasoning. Our work establishes a promising direction for data-efficient video-language modeling.

视频语言模型时间推理多模态融合数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。