arXiv:2411.02851cs.MMcs.AI2024-11被引 3

融合音视频文本三模态,提升多语言视频问答定位准确率

Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization

  • 构建三模态融合框架,分别设计音视、视觉、文本预测器
  • 引入动态三角损失函数,实现跨模态协同学习与结果一致
  • 在MVAL任务上超越多个SOTA方法,验证音频模态价值

多语言视频问答定位(MVAL)旨在定位回答给定多语言问题的视频片段。现有方法或仅依赖视觉模态,或融合视觉与字幕模态,但忽略了视频中的音频信息,导致输入不完整,影响性能。本文提出统一的音视频文本跨度定位(AVTSL)方法,引入音频模态以增强视觉与文本表征。具体地,融合三模态特征并设计三个专属预测器:音视预测器、视觉预测器和文本预测器,各基于自身模态生成预测。为保证预测一致性,提出音视频文本一致性模块,采用动态三角损失(DTL)函数,使各模态预测器可动态学习其他模态输出。实验表明,该方法显著优于多个SOTA模型,验证了音频模态的有效性。

原文摘要 · Abstract (English)

The goal of Multilingual Visual Answer Localization (MVAL) is to locate a video segment that answers a given multilingual question. Existing methods either focus solely on visual modality or integrate visual and subtitle modalities. However, these methods neglect the audio modality in videos, consequently leading to incomplete input information and poor performance in the MVAL task. In this paper, we propose a unified Audio-Visual-Textual Span Localization (AVTSL) method that incorporates audio modality to augment both visual and textual representations for the MVAL task. Specifically, we integrate features from three modalities and develop three predictors, each tailored to the unique contributions of the fused modalities: an audio-visual predictor, a visual predictor, and a textual predictor. Each predictor generates predictions based on its respective modality. To maintain consistency across the predicted results, we introduce an Audio-Visual-Textual Consistency module. This module utilizes a Dynamic Triangular Loss (DTL) function, allowing each modality's predictor to dynamically learn from the others. This collaborative learning ensures that the model generates consistent and comprehensive answers. Extensive experiments show that our proposed method outperforms several state-of-the-art (SOTA) methods, which demonstrates the effectiveness of the audio modality.

多模态视频问答音频增强跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。