arXiv:2605.12036eess.AS2026-05

构建14维语音理解基准,提升模型对细微语音特征的感知能力

Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model

论文配图:Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
图 1 · 摘自论文原文
  • 构建细粒度多维语音理解数据集与评估框架
  • 提出FM-Speech模型,在14维属性上显著超越现有开源模型
  • 解决真实场景下语音细节感知难题,适合语音系统研发者

尽管语音大语言模型在语音识别等常规任务中表现优异,但在微小声学线索、声学场景和副语言信号等细粒度、多维度特征的分离理解上仍存在明显不足。这一感知缺陷导致对真实语音的不完整理解,制约了下一代感知型、共情型语音系统的发展。根源在于高质量表达数据稀缺、多维属性建模缺失以及评测覆盖范围有限、粒度粗略。本文从三方面应对:首先,建立稳健的数据清洗流程,从视听资源中提取高质量自发语音语料,解决复杂声学环境与长音频时间对齐问题;其次,构建首个涵盖14个语音属性维度的基准FMSU-Bench,用于严格评估模型的细粒度多维语音理解能力;第三,基于所获语料,提出FM-Speech模型,采用解耦属性建模与渐进式课程微调框架,显著提升细粒度多维声学感知能力。大量实验表明,当前语音大模型在多维细粒度理解上仍有较大提升空间,而FM-Speech显著优于现有开源模型,为真实世界语音理解建立了新范式。

原文摘要 · Abstract (English)

While speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like micro-acoustic cues, acoustic scenes, and paralinguistic signals. This resulting incomplete comprehension of real-world speech fundamentally bottlenecks the development of perceptive and empathetic next-generation speech systems. At its core, this persistent perceptual limitation primarily stems from three interacting factors: scarce high-quality expressive data, absent fine-grained modeling for multi-dimensional attributes, and reliance on restricted coverage, coarse-grained benchmarks. We address these challenges through three pillars: First, our robust data curation pipeline resolves complex acoustic environments and long-audio timestamp alignment challenges to extract a high-quality spontaneous speech corpus from audiovisual sources. Second, we construct FMSU-Bench, a pioneering benchmark covering 14 speech attribute dimensions to rigorously assess the fine-grained, multi-dimensional speech understanding capabilities of current models. Third, empowered by our curated corpus, we introduce FM-Speech. Driven by a decoupled attribute modeling and progressive curriculum fine-tuning framework, it substantially elevates fine-grained, multi-dimensional acoustic perception. Extensive evaluations on FMSU-Bench reveal that current speech LLMs still require significant improvement in multi-dimensional, fine-grained understanding. In contrast, FM-Speech substantially outperforms current open-source models, establishing a robust paradigm for real-world speech understanding.

语音理解多维感知大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。