构建音频嵌入评估基准,推动多模态系统听觉智能发展。
Massive Sound Embedding Benchmark (MSEB)
- 提出可扩展的音频嵌入评测框架MSEB,涵盖8项核心任务。
- 基于新数据集SVQ等,首次实验揭示显著性能差距。
- 适合研究多模态感知、音频理解与模型评估的学者使用。
音频是多模态感知的关键组成部分,任何真正智能的系统都必须具备广泛的听觉能力,包括转录、分类、检索、推理、分割、聚类、重排序和重建。这些任务本质上都涉及将原始音频信号转换为有意义的‘嵌入’——无论是单个向量、连续或离散表示序列,或其他结构化形式——作为生成最终响应的基础。为加速实现稳健的机器听觉智能,我们提出了大规模音频嵌入基准(Massive Sound Embedding Benchmark, MSEB):一个可扩展的框架,用于评估任意多模态系统的听觉组件。在首版中,MSEB 提供了8项核心任务,涵盖多样化数据集,包括新发布的大型简单语音问题(Simple Voice Questions, SVQ)数据集。初步实验揭示了明显的性能提升空间,凸显了在以音频为核心信号的真实多模态体验中改进的巨大潜力。我们鼓励研究社区使用 MSEB 评估算法并参与其持续发展。相关库已公开于 GitHub。
原文摘要 · Abstract (English)
Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding' - be it a single vector, a sequence of continuous or discrete representations, or another structured form - which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。