arXiv:2605.03398cs.CV2026-05

用大模型辅助对齐视频语义与时间,提升定位准确性。

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

论文配图:MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
图 1 · 摘自论文原文
  • 训练时用大模型生成事件级描述和片段级标题,指导模型对齐语义与时间。
  • 通过增强事件语义对齐和局部关系一致性,显著提升时间跨度区分度。
  • 适合需要精准视频时间定位的研究者,尤其在复杂场景下表现更优。

视频时间定位(VTG)面临跨模态语义鸿沟问题,常导致背景特征错误对齐查询,而直接匹配查询与时间段又存在判别力不足和时间语义不一致。为此,我们提出一种基于多模态大模型(MLLM)的训练阶段优化框架MASRA。MASRA利用MLLM生成两类文本先验:带时间跨度的事件级描述和片段级标题,并构建两种对齐机制:事件语义时间对齐(ESTA)通过强化语义与时间事件的对应关系,提升段落级可分性;局部关系一致性对齐(LRCA)基于片段标题构建文本关系矩阵,与模型中的时间特征相似矩阵对齐,增强时间一致性并捕捉局部结构信息。MASRA还引入语义引导增强与二阶关系注意力两个模块,更好利用学习到的语义上下文与关系结构。此外,提出解耦对齐交互(DAI)结合上下文感知码本,自适应过滤无关语义,缓解跨模态差异。MLLM仅用于训练,推理时无需参与。大量实验表明MASRA优于现有方法,消融实验证实其有效性。

原文摘要 · Abstract (English)

Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query to moments results in insufficient discriminability and consistency of temporal semantics. To address this issue, we propose MLLM-Assisted Semantic-Relational Consistent Alignment (MASRA), a training-time MLLM-based optimization framework for VTG. MASRA leverages an MLLM during training to produce two forms of textual priors, namely event-level descriptions with temporal spans and clip-level captions, and instantiates two MLLM-assisted alignments. Event Semantic Temporal Alignment (ESTA) aligns temporal context with event semantics to explicitly strengthen the correspondence between semantics and temporal events and improve span-level separability. Local Relational Consistency Alignment (LRCA) constructs a textual relation matrix derived from clip-level captions and aligns it with the temporal feature similarity matrix in the model, enhancing temporal consistency while capturing local structural information. MASRA includes two simple supporting modules, semantic-guided enhancement and second-order relational attention, to better utilize the learned semantic context and relational structure. Moreover, we introduce Decoupled Alignment Interaction (DAI) with a context-aware codebook to adaptively absorb query-irrelevant semantics and alleviate the cross-modal gap. The MLLM is only invoked during training and is not used at inference. Extensive experiments show that MASRA outperforms existing methods, and ablation studies validate its effectiveness.

视频定位多模态大模型时间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。