基于EnCLAP框架优化音频描述生成,提升自动音频字幕准确率。
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
- 在EnCLAP基础上改进组件并引入重排序机制
- 音频字幕任务取得0.542的FENSE得分
- 衍生出检索模型,适用于语言驱动音频搜索
本文描述了我们针对DCASE2024挑战赛任务6(自动音频字幕)和任务8(基于语言的音频检索)的提交方案。我们的方法在EnCLAP音频字幕框架基础上进行改进,并针对任务6进行了优化。文中详细说明了底层组件的调整及重排序流程的引入。此外,我们还提交了一个辅助检索模型,该模型是改进后框架的副产品,用于任务8。所提出的系统在任务6上获得0.542的FENSE得分,在任务8上取得0.386的mAP@10得分,显著优于基线模型。
原文摘要 · Abstract (English)
In this technical report, we describe our submission to DCASE2024 Challenge Task6 (Automated Audio Captioning) and Task8 (Language-based Audio Retrieval). We develop our approach building upon the EnCLAP audio captioning framework and optimizing it for Task6 of the challenge. Notably, we outline the changes in the underlying components and the incorporation of the reranking process. Additionally, we submit a supplementary retriever model, a byproduct of our modified framework, to Task8. Our proposed systems achieve FENSE score of 0.542 on Task6 and mAP@10 score of 0.386 on Task8, significantly outperforming the baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。