用最优传输融合音视频,提升自动音频描述生成质量
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
- 通过最优传输对齐音视频特征,解决模态间语义鸿沟
- 在AudioCaps上超越现有方法,无需大规模数据或后处理
- 适合多模态内容理解与智能生成任务的研究者
自动音频字幕生成旨在为音频内容生成文本描述,近期研究尝试利用视觉信息提升生成质量。然而,现有方法往往难以有效融合音视频数据,忽略各模态的重要语义线索。为此,我们提出LAVCap,一个基于大语言模型的音视频字幕生成框架,通过最优传输对齐损失有效整合视觉与音频信息,提升字幕生成性能。此外,我们设计了最优传输注意力模块,利用最优传输分配图增强音视频融合。结合优化训练策略,实验表明框架各组件均有效。LAVCap在AudioCaps数据集上优于现有最先进方法,且无需依赖大规模数据集或后处理。代码已开源。
原文摘要 · Abstract (English)
Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from each modality. To address this, we introduce LAVCap, a large language model (LLM)-based audio-visual captioning framework that effectively integrates visual information with audio to improve audio captioning performance. LAVCap employs an optimal transport-based alignment loss to bridge the modality gap between audio and visual features, enabling more effective semantic extraction. Additionally, we propose an optimal transport attention module that enhances audio-visual fusion using an optimal transport assignment map. Combined with the optimal training strategy, experimental results demonstrate that each component of our framework is effective. LAVCap outperforms existing state-of-the-art methods on the AudioCaps dataset, without relying on large datasets or post-processing. Code is available at https://github.com/NAVER-INTEL-Co-Lab/gaudi-lavcap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。