arXiv:2509.07526cs.SDcs.AI2025-09中稿 · ASRU 2025被引 3

用不到3万小时公开音频数据,训练出高效且性能顶尖的音视频模型。

Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data

  • 基于指令微调大模型与Whisper编码器,单阶段训练实现音视频理解。
  • 7B模型在MMAU上达64.14分,超越多数开源大模型,仅需5000小时独特数据。
  • 小至1B参数模型仍媲美2-13B大模型,适合资源有限的研究者使用。

大语言模型(LLMs)已彻底改变自然语言处理,但其与音频的融合仍进展缓慢,尽管音频是人类交流的核心。我们提出Falcon3-Audio,一类基于指令微调的大型语言模型与Whisper编码器构建的音视频模型(ALMs)。仅使用少于3万小时(5000个独特)的公共音频数据,Falcon3-Audio-7B在MMAU基准上取得64.14分,达到当前开源权重模型最佳水平,与R1-AQA持平。该模型在数据效率、参数效率、训练流程透明性方面表现卓越。值得注意的是,其最小的1B模型仍能与2-13B参数的开源大模型竞争。通过大量消融实验发现,复杂的教学策略、多音频编码器或复杂交叉注意力结构并非强性能必要条件,甚至优于训练于超50万小时数据的模型。

原文摘要 · Abstract (English)

Large language models (LLMs) have transformed NLP, yet their integration with audio remains underexplored despite audio's centrality to human communication. We introduce Falcon3-Audio, a family of Audio-Language Models (ALMs) built on instruction-tuned LLMs and Whisper encoders. Using a remarkably small amount of public audio data, less than 30K hours (5K unique), Falcon3-Audio-7B matches the best reported performance among open-weight models on the MMAU benchmark, with a score of 64.14, matching R1-AQA, while distinguishing itself through superior data and parameter efficiency, single-stage training, and transparency. Notably, our smallest 1B model remains competitive with larger open models ranging from 2B to 13B parameters. Through extensive ablations, we find that common complexities such as curriculum learning, multiple audio encoders, and intricate cross-attention connectors are not required for strong performance, even compared to models trained on over 500K hours of data.

音视频模型数据效率单阶段训练Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。