用生成的文本描述增强音频检索,提升语音字幕生成效果
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
- 用自动生成的文本描述作为多模态查询,匹配知识库中的音视频对
- 在 AudioCaps、Clotho 等数据集上达到当前最优性能
- 渐进式学习策略让模型逐步融合更多音视频配对,训练更稳定
检索增强生成可通过从知识库中引入相关音视频对来提升语音字幕生成质量。现有方法通常仅使用输入音频作为单模态查询。本文提出生成辅助的多模态查询方法,通过生成输入音频的文本描述,使查询模态与知识库中的音视频结构一致,从而实现更有效的检索。此外,我们设计了一种新的渐进式学习策略,逐步增加交错音视频对的数量以优化训练过程。在 AudioCaps、Clotho 以及 Auto-ACD 数据集上的实验表明,该方法在多个基准上均达到当前最优性能。
原文摘要 · Abstract (English)
Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose Generation-Assisted Multimodal Querying, which generates a text description of the input audio to enable multimodal querying. This approach aligns the query modality with the audio-text structure of the knowledge base, leading to more effective retrieval. Furthermore, we introduce a novel progressive learning strategy that gradually increases the number of interleaved audio-text pairs to enhance the training process. Our experiments on AudioCaps, Clotho, and Auto-ACD demonstrate that our approach achieves state-of-the-art results across these benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。