用有监督方式生成音频标记,提升自动音频描述性能。
Discrete Audio Representations for Automated Audio Captioning
- 提出有监督音频标记方法,基于音频分类目标训练。
- 在Clotho数据集上,新标记使自动音频描述准确率更高。
- 相比无监督标记,新方法更懂音频事件含义,适合语音理解任务。
离散音频表示(称为音频标记)分为语义和声学两类,通常通过连续音频表示的无监督分词生成。然而,其在自动音频描述(AAC)中的适用性仍不明确。本文通过对比多种分词方法,系统评估了音频标记驱动模型在AAC中的可行性。结果表明,与直接使用连续音频表示的模型相比,音频标记化会导致性能下降。为解决此问题,我们提出一种基于音频标注目标的有监督音频分词器。不同于缺乏显式语义理解的无监督分词器,该分词器能有效捕捉音频事件信息。在Clotho数据集上的实验显示,所提出的音频标记在自动音频描述任务中优于传统音频标记。
原文摘要 · Abstract (English)
Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。