arXiv:2507.14129cs.SDeess.AS2025-07中稿 · WASPAA 2025被引 8

开源音频编码器OpenBEATs用多领域数据提升音频理解能力,性能超越大模型。

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

  • 基于掩码令牌预测,在多领域音频上预训练通用音频表示。
  • 在6个生物声学、2个环境声音和5个推理任务上达顶尖性能。
  • 模型参数仅为大模型1/4,但效果更优,适合音频研究与应用者使用。

掩码令牌预测已成为语言、视觉和语音领域的强大预训练目标,有望通过单一预训练任务统一这些异构模态。然而,其在通用音频理解中的应用仍不充分,仅有BEATs是显著案例。由于缺乏开源预训练代码,BEATs的改进有限,且仅在AudioSet上训练,限制了其下游泛化能力。为此,我们提出OpenBEATs,一个开源框架,通过多领域音频预训练扩展BEATs。我们在六类任务、二十五个数据集、三个音频领域上进行综合评估,涵盖音频问答、蕴含关系和图像描述等推理任务。OpenBEATs在六个生物声学数据集、两个环境声音数据集及五个推理数据集上达到当前最优表现,优于参数量超十亿的模型,且仅为其四分之一。结果证明,多领域数据与掩码令牌预测任务对学习通用音频表征的有效性。为促进研究与可复现性,我们公开所有预训练与评估代码、预训练及微调检查点和训练日志,详见 https://github.com/Shikhar-S/OpenBEATs

原文摘要 · Abstract (English)

Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general audio understanding remains underexplored, with BEATs being the only notable example. BEATs has seen limited modifications due to the absence of open-source pre-training code. Furthermore, BEATs was trained only on AudioSet, restricting its broader downstream applicability. To address these gaps, we present OpenBEATs, an open-source framework that extends BEATs via multi-domain audio pre-training. We conduct comprehensive evaluations across six types of tasks, twenty five datasets, and three audio domains, including audio reasoning tasks such as audio question answering, entailment, and captioning. OpenBEATs achieves state-of-the-art performance on six bioacoustics datasets, two environmental sound datasets and five reasoning datasets, performing better than models exceeding a billion parameters at one-fourth their parameter size. These results demonstrate the effectiveness of multi-domain datasets and masked token prediction task to learn general-purpose audio representations. To promote further research and reproducibility, we release all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs at https://github.com/Shikhar-S/OpenBEATs

音频编码多模态开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。