统一语音理解实验框架,提升模型可比性与复现性。
A Unified and Reproducible Experimentation Framework for Speech Understanding

- 标准化预测格式、归一化和评分方式,统一评估流程。
- 在真实声学与语言压力下,跨传统流水线与语音大模型评估性能。
- 引入智能代理辅助转换训练流程,支持开源数据集上的可复现训练。
语音基础模型和语音大模型推动了语音理解的发展,但部署导向的模型选择因后处理不一致导致评估不可比,且不同数据规模和流水线下的训练结果难以复现。我们提出SURE框架,统一预测格式、归一化和评分标准。SURE在真实声学与语言压力下,对从传统流水线到语音大模型的多种系统进行代表性任务评估。除评估外,SURE还引入代理辅助的训练转换流程,将论文与代码映射为版本化、可运行的训练流水线,基于匹配的开源数据子集。整体上,SURE提升了部署导向评估的可比性与可复现性。
原文摘要 · Abstract (English)
Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched post-processing, and by training results that are hard to reproduce across data scales and pipelines. We present SURE, a unified experimentation framework that standardizes prediction formats, normalization, and scoring. SURE evaluates strong systems across paradigms, from conventional pipelines to Speech LLMs, on representative tasks under realistic acoustic and linguistic stressors. Beyond evaluation, SURE introduces an agent-assisted training conversion flow that maps paper and code into versioned, runnable training pipelines under a unified protocol on matched open-data subsets. Overall, SURE improves comparability and reproducibility for deployment-oriented evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。