arXiv:2607.01667cs.CV2026-07

解决音视频描述中时间与模态对齐难题,提升叙事连贯性。

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning

论文配图:Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning
图 1 · 摘自论文原文
  • 采用观察-检验-修正迭代框架生成高质量训练数据
  • 在真实人类交互数据集上实现精准音视频绑定与时序推理
  • 构建解耦评估基准,可量化分析模型对齐能力

尽管多模态大语言模型已推动视频理解发展,但在音视频视频描述中实现精确的时间与跨模态对齐仍是重大挑战。现有方法普遍存在模态脱节和时序不一致问题,难以准确关联听觉事件与视觉实体,也难以捕捉复杂的因果动态。为此,我们提出TCA-Captioner,一个专为增强音视频描述中的时间与跨模态对齐而设计的框架。首先引入观察-检验-修正(OCC)框架,通过迭代优化生成高保真、精确对齐的训练数据。基于精心筛选的高密度人类交互数据集,TCA-Captioner被优化以建模复杂的音视频交互关系。此外,我们提出TCA-Bench诊断基准,采用解耦评估协议,独立量化模型在音视频绑定与时间关系推理方面的表现。大量实验表明,TCA-Captioner在时间连贯性和音视频同步叙述方面树立了新标准。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have advanced video understanding, achieving precise temporal and cross-modal alignment in audiovisual video captioning remains a formidable challenge. Most existing approaches suffer from modality detachment and temporal incoherence, failing to accurately bind auditory events to visual entities or capture complex causal dynamics. To address these deficiencies, we propose TCA-Captioner, a framework specifically engineered to enhance Temporal and Cross-Modal Alignment for audiovisual video captioning. We first introduce the Observer-Checker-Corrector (OCC) framework, an iterative refinement strategy that generates high-fidelity, meticulously grounded training data. Leveraging a curated high-density human interaction dataset, TCA-Captioner is optimized to model sophisticated audiovisual interactions. Furthermore, we present TCA-Bench, a diagnostic benchmark utilizing a Decoupled Evaluation Protocol to isolate and quantify model proficiency in audiovisual binding and temporal relational reasoning. Extensive experiments demonstrate that TCA-Captioner sets a new standard for temporally-coherent and synchronized audiovisual narratives.

音视频对齐视频描述多模态时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。