对比两种音频预训练方法,发现各有优劣,为通用音频表征学习指明方向。
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- 构建1070万条跨领域音频-文本数据集CaptionStew,解决数据不足问题。
- 对比对比学习与描述生成任务,发现前者更省数据,后者更易扩展。
- 揭示大模型下监督初始化优势减弱,挑战现有训练范式。
音频-语言预训练(ALP)有望学习通用音频表征,但研究仍不充分。当前缺乏共识:音频-语言模型能否构建有效的通用音频编码器,以及预训练目标在不同任务与规模下的表现机制尚不清晰。我们识别出三大障碍:音频-文本语料库规模有限、现有字幕数据集覆盖的音频属性不足、系统性探索与评估缺失。为此,我们首次开展系统性实证研究。首先提出CaptionStew,一个包含1070万条字幕的多领域、多关注点开源音频-文本数据集。随后,首次全面比较对比学习与字幕生成目标在语音、音乐和环境音任务上的表现。结果表明,ALP可获得具有竞争力且可迁移的音频表征,同时揭示关键权衡:对比学习具备更优的数据效率,而字幕生成具备更好可扩展性。此外,我们发现监督初始化的优势在更大规模下常会减弱,挑战了常见做法。基于实证证据,我们建立了一条通往通用音频表征学习的可行路径,为未来研究提供指导。
原文摘要 · Abstract (English)
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-language models can build effective general-purpose audio encoders, nor a systematic understanding of how pretraining objectives behave across diverse tasks and scales. We identify three key barriers: limited scale of audio-text corpora, limited coverage of audio attributes in existing caption corpora, and lack of systematic exploration and evaluation. To fill this gap, we present the first principled empirical study of ALP. We first introduce CaptionStew, a 10.7M caption dataset aggregating open-source audio-text corpora across multiple domains and captioning focuses. We then conduct the first comprehensive evaluation comparing contrastive and captioning objectives for learning audio representation across speech, music, and environmental sound tasks. Our results not only demonstrate that ALP yields competitive, transferable representations, but reveal critical trade-offs: contrastive learning offers superior data efficiency, while captioning exhibits better scalability. Furthermore, we find that the benefits of supervised initialization often diminish at larger scales, challenging common practices. By grounding these claims in empirical evidence, we establish a viable pathway toward general-purpose audio representation learning, guiding future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。