为不同语音模型匹配合适评估方法,解决评价混乱问题
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
- 提出三维评估分类框架,按能力、任务和评价维度划分
- 梳理现有评测体系,发现韵律、交互、推理等覆盖不足
- 帮助研究者选对评估方式,推动未来评测设计改进
语音基础模型在多项任务中表现出色,但其评估方法分散且不统一。不同模型擅长语音处理的不同方面,需采用不同评估协议。本文提出一个统一的分类框架,回答‘哪种模型应配哪种评估’的问题。该框架包含三个正交维度:评估所关注的方面、模型完成任务所需的能力,以及执行任务所需的协议要求。我们沿这三个维度对广泛存在的评估与基准进行了分类,涵盖表征学习、语音生成和交互对话等领域。通过将每项评估映射到模型暴露的能力(如语音生成、实时处理)及其方法学需求(如微调数据、人工判断),该框架为模型与评估方法的匹配提供了系统性指导。同时揭示出当前评估体系在韵律、交互和推理等方面的系统性缺失,指明未来基准设计的关键方向。本工作为语音模型评估的选择、解读与拓展提供了概念基础与实践指南。
原文摘要 · Abstract (English)
Speech foundation models have recently achieved remarkable capabilities across a wide range of tasks. However, their evaluation remains disjointed across tasks and model types. Different models excel at distinct aspects of speech processing and thus require different evaluation protocols. This paper proposes a unified taxonomy that addresses the question: Which evaluation is appropriate for which model? The taxonomy defines three orthogonal axes: the evaluation aspect being measured, the model capabilities required to attempt the task, and the task or protocol requirements needed to perform it. We classify a broad set of existing evaluations and benchmarks along these axes, spanning areas such as representation learning, speech generation, and interactive dialogue. By mapping each evaluation to the capabilities a model exposes (e.g., speech generation, real-time processing) and to its methodological demands (e.g., fine-tuning data, human judgment), the taxonomy provides a principled framework for aligning models with suitable evaluation methods. It also reveals systematic gaps, such as limited coverage of prosody, interaction, or reasoning, that highlight priorities for future benchmark design. Overall, this work offers a conceptual foundation and practical guide for selecting, interpreting, and extending evaluations of speech models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。