评测主流大模型在八大音频能力上的表现,发现音视频差距仍存。
Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

- 用MSEB基准测试GPT、Gemini等大模型的音频理解能力
- 模型在八项核心任务中表现不一,音视频性能仍有明显差距
- 适合关注多模态模型选型与应用落地的研究者
大规模声音嵌入基准(MSEB)已成为评估音频模型功能广度的标准。尽管早期基线聚焦于专用编码器,但向‘音频原生’大型语言模型(LLMs)的转变暗示了一种新范式:单一多模态主干网络或可替代复杂的任务专用流水线。本文对领先的大模型(包括Gemini和GPT系列)在MSEB的八个核心能力上进行了严格的实证评估,以检验其有效性及音视频对齐程度。结果表明,尽管性能与鲁棒性方面仍存在显著模态差距,但支持‘最优’建模方法的实证证据尚不充分。最终,音频原生与级联架构的选择高度依赖具体应用场景,以及对延迟、成本和推理深度的假设。
原文摘要 · Abstract (English)
The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines focused on specialized encoders, the shift toward "audio-native" Large Language Models (LLMs) suggests a new paradigm where a single multimodal backbone may replace complex, task-specific pipelines. This paper provides a rigorous empirical evaluation of leading LLMs - including members from the Gemini and GPT families - across the eight core MSEB capabilities to assess their efficacy and audio-text parity. Our results indicate that while a significant modality gap persists regarding performance and robustness, the empirical evidence for an "optimal" modeling approach remains inconclusive. Ultimately, the choice between audionative and cascaded architectures depends heavily on specific use-case requirements and the underlying assumptions regarding latency, cost, and reasoning depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。