用语义丰富的离散音频标记提升自动音频描述效果
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
- 通过向量量化从预训练音频表示中生成语义丰富的离散标记
- 在两个基准上优于基线EnCLAP,性能显著提升
- 适合需要精准音频语义理解的研究者与开发者
自动音频描述(AAC)旨在通过有效的声学特征描述通用声音的语义上下文,包括声音事件和场景。为提升性能,基线方法EnCLAP采用EnCodec生成的离散标记作为语言模型BART微调的输入。然而,EnCodec的设计目标是波形重建而非捕捉通用声音的语义上下文,这限制了其在AAC中的表现。为此,我们提出CLAP-ART,一种利用“语义丰富且离散”标记的AAC方法。CLAP-ART通过向量量化从预训练音频表示(AR)中计算出语义丰富的离散标记。实验表明,CLAP-ART在两个AAC基准上均优于基线EnCLAP,证明了从语义丰富的音频表示中提取的离散标记对自动音频描述具有显著促进作用。
原文摘要 · Abstract (English)
Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhance performance, an AAC method, EnCLAP, employed discrete tokens from EnCodec as an effective input for fine-tuning a language model BART. However, EnCodec is designed to reconstruct waveforms rather than capture the semantic contexts of general sounds, which AAC should describe. To address this issue, we propose CLAP-ART, an AAC method that utilizes ``semantic-rich and discrete'' tokens as input. CLAP-ART computes semantic-rich discrete tokens from pre-trained audio representations through vector quantization. We experimentally confirmed that CLAP-ART outperforms baseline EnCLAP on two AAC benchmarks, indicating that semantic-rich discrete tokens derived from semantically rich AR are beneficial for AAC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。