arXiv:2601.05495cs.CVcs.CL2026-01被引 2

为长视频设计多粒度表示,提升理解效率与精度

MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding

  • 通过关键转折点分段,构建三层结构化表示
  • 在小时级视频理解上提升19.67%,延迟降至45.4%
  • 适合需要高效长视频分析的场景,如智能检索

时长从几分钟到数小时的长视频,因其复杂事件、多样场景和长程依赖,对当前多模态大模型(MLLMs)构成重大挑战。直接编码此类视频计算开销过大,而简单视频转文本常导致内容冗余或碎片化。为此,我们提出MMViR——一种新型多模态、多粒度结构化表示方法,用于长视频理解。MMViR识别关键转折点进行视频分段,并构建耦合全局叙事与细粒度视觉细节的三层描述体系。该设计支持高效查询式检索,且在多种场景下具有强泛化能力。在问答、摘要和检索三项任务上的广泛评估表明,MMViR优于现有最强方法,在小时级视频理解上实现19.67%的性能提升,同时将处理延迟降低至原始水平的45.4%。

原文摘要 · Abstract (English)

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is computationally too expensive, while simple video-to-text conversion often results in redundant or fragmented content. To address these limitations, we introduce MMViR, a novel multi-modal, multi-grained structured representation for long video understanding. MMViR identifies key turning points to segment the video and constructs a three-level description that couples global narratives with fine-grained visual details. This design supports efficient query-based retrieval and generalizes well across various scenarios. Extensive evaluations across three tasks, including QA, summarization, and retrieval, show that MMViR outperforms the prior strongest method, achieving a 19.67% improvement in hour-long video understanding while reducing processing latency to 45.4% of the original.

长视频理解多模态结构化表示视频检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。