arXiv:2411.06872cs.CVcs.AI2024-11被引 2

融合音视频信息生成可解释的精准视频描述

Multi-Modal interpretable automatic video captioning

  • 采用多模态对比学习,融合视觉与音频特征
  • 在MSR-VTT和VATEX数据集上优于现有模型
  • 通过注意力机制提供决策过程解释,适合需要透明性的应用

视频字幕任务旨在用自然语言描述视频内容,需理解场景、动作与事件的同步发生。当前方法多聚焦视觉线索,常忽略音频等重要模态及其相互依赖关系。本文提出一种基于多模态对比损失的新方法,强调多模态融合与可解释性。该方法能捕捉模态间依赖关系,生成更准确、相关的字幕。同时引入多重注意力机制,揭示模型决策过程。实验表明,该方法在常用基准数据集MSR-VTT和VATEX上优于现有最先进模型。

原文摘要 · Abstract (English)

Video captioning aims to describe video contents using natural language format that involves understanding and interpreting scenes, actions and events that occurs simultaneously on the view. Current approaches have mainly concentrated on visual cues, often neglecting the rich information available from other important modality of audio information, including their inter-dependencies. In this work, we introduce a novel video captioning method trained with multi-modal contrastive loss that emphasizes both multi-modal integration and interpretability. Our approach is designed to capture the dependency between these modalities, resulting in more accurate, thus pertinent captions. Furthermore, we highlight the importance of interpretability, employing multiple attention mechanisms that provide explanation into the model's decision-making process. Our experimental results demonstrate that our proposed method performs favorably against the state-of the-art models on commonly used benchmark datasets of MSR-VTT and VATEX.

视频生成多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。