首个评测语音模型发音识别能力的开源基准,揭示现有系统盲点。
PRiSM: Benchmarking Phone Realization in Speech Models
- 构建多维度评测框架,结合内在与外在评估方式。
- 发现训练时接触多种语言显著提升发音识别性能。
- 适合关注跨语言语音模型、语音分析的研究者使用。
发音识别(PR)是实现无语言依赖的跨语言语音处理与音位分析的基础接口。尽管长期致力于发展发音识别系统,当前评估仍仅关注表面转写准确率。我们提出PRiSM,首个开源基准,旨在通过内在与外在评估揭示发音感知中的盲点。PRiSM统一了基于转写的评估标准,并通过转写与表征探针,在临床、教育及多语言场景中评估下游实用性。实验发现,训练时多样化语言暴露对发音识别性能至关重要,编码器-CTC模型表现最稳定,且专用发音识别模型仍优于大型音频语言模型。PRiSM发布代码、训练配方与数据集,推动具备稳健发音能力的多语言语音模型发展:https://github.com/changelinglab/prism。
原文摘要 · Abstract (English)
Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. We find that diverse language exposure during training is key to PR performance, encoder-CTC models are the most stable, and specialized PR models still outperform Large Audio Language Models. PRiSM releases code, recipes, and datasets to move the field toward multilingual speech models with robust phonetic ability: https://github.com/changelinglab/prism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。