arXiv:2603.25767cs.SDcs.AI2026-03中稿 · CVPR被引 1

构建强监督数据体系,提升音频预训练质量

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

  • 用高保真字幕生成器和统一标签系统构建强标注数据
  • 高质量数据使模型性能显著优于传统弱标签方法
  • 适合音频理解、多模态学习研究者参考

当前音频预训练旨在为广泛音频理解任务学习统一表征,但仍因依赖弱标签、噪声大且规模有限而发展受限。受视觉领域基础预训练范式的启发,我们主张音频领域应首先建立自身的大规模强监督框架。本文提出一种新的数据驱动流程:利用高保真字幕生成器创建业界领先水平的字幕,并首次构建统一标签系统(UTS),实现语音、音乐与环境音的跨模态对齐。在此强监督数据基础上,系统比较不同预训练目标的效果。实验表明,数据质量和覆盖范围是性能提升的核心驱动力,而预训练目标则决定下游任务的专属性能。

原文摘要 · Abstract (English)

Current audio pre-training seeks to learn unified representations for broad audio understanding tasks, but it remains fragmented and is fundamentally bottlenecked by its reliance on weak, noisy, and scale-limited labels. Drawing lessons from vision's foundational pre-training blueprint, we argue that the audio field must first establish its own large-scale, strong supervision framework. We introduce a new data-centric pipeline that leverages a high-fidelity captioner to create SOTA-quality captions and the first Unified Tag System (UTS) that bridges speech, music, and environmental sounds. We then conduct a systematic comparative study of different pre-training objectives on these strong source data. Our experiments suggest that data quality and coverage are the primary drivers of performance, while the choice of objective dictates downstream task specialization.

音频预训练强监督数据构建统一标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。