1.7B参数模型实现顶尖音频理解能力,体积小但性能强。
Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
- 用轻量语言主干+Whisper音频编码器+稀疏专家模块,统一建模音频
- 1.7B参数模型在多项任务上超越7B至30B大模型
- 适合资源有限场景下的高效音频智能应用
我们提出Eureka-Audio,一个仅含1.7B参数的紧凑型音频语言模型,在广泛的音频理解基准测试中表现媲美规模大4至18倍的模型。该模型采用统一端到端架构,包含轻量语言主干、基于Whisper的音频编码器和稀疏激活的混合专家(MoE)适配器,有效应对音频异质性并缓解跨模态优化冲突。为增强副语言推理能力,我们设计DataFlux——一种闭环音频指令数据合成与验证流水线,从原始音频生成高质量、逻辑一致的监督信号。在自动语音识别(ASR)、知识推理、安全、指令遵循及副语言基准上的全面评估表明,Eureka-Audio在计算成本与性能间实现了高效平衡。这些结果确立了Eureka-Audio作为轻量级音频理解模型的强有力实用基准。
原文摘要 · Abstract (English)
We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Despite containing only 1.7B parameters, Eureka-Audio demonstrates strong performance on automatic speech recognition (ASR), audio understanding, and dense audio captioning, matching or surpassing multiple 7B to 30B audio and omni-modal baselines. The model adopts a unified end-to-end architecture composed of a lightweight language backbone, a Whisper-based audio encoder, and a sparsely activated Mixture-of-Experts (MoE) adapter that explicitly accounts for audio heterogeneity and alleviates cross-modal optimization conflicts under limited capacity. To further enhance paralinguistic reasoning, we introduce DataFlux, a closed loop audio instruction data synthesis and verification pipeline that constructs high quality, logically consistent supervision from raw audio. Extensive evaluations across ASR, knowledge reasoning, safety, instruction following, and paralinguistic benchmarks, demonstrate that Eureka-Audio achieves an efficient balance between computational cost and performance. These results establish Eureka Audio as a strong and practical baseline for lightweight audio understanding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。