arXiv:2606.18466cs.CL2026-06

MFA 3.0实现语音对齐新基准,支持多语言且误差低于15毫秒

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

  • 基于大规模语料和跨语言映射,提升多语言对齐能力
  • 在英日韩三语上均达最先进水平,平均边界误差<15毫秒
  • 适合需要高精度语音对齐的语音研究与工业应用

蒙特利尔强制对齐器(MFA)自2016年发布以来,已成为科研与工业界最广泛使用的强制对齐工具。十年间,MFA通过使用更大规模的开源数据集,扩展了对更多语言和方言的支持,统一了国际音标词典,引入模型适应、跨语言音素映射及辅助工具等功能。本文记录了MFA 3.0相较于1.0版本的主要进展,并在英语、日语和韩语上评估其性能,对比经典与神经网络对齐器。MFA 3.0在四个基准数据集上均达到或接近当前最优水平,平均边界误差低于15毫秒。模型适应与跨语言音素重映射对训练分布外语言有效;发音概率建模与音系规则在特定条件下带来性能提升。

原文摘要 · Abstract (English)

The Montreal Forced Aligner (MFA) was released in 2016 and has since become the most widely used tool for forced alignment in research and industry. In the decade since, MFA has undergone substantial development, including expanded coverage across more languages and dialects using larger open-source datasets, harmonized IPA dictionaries, model adaptation, cross-language phone remapping, and support utilities. This paper documents MFA 3.0's developments since version 1.0 and evaluates MFA's performance across English, Japanese, and Korean, benchmarked against classic and neural forced aligners. MFA 3.0 achieves state-of-the-art or near state-of-the-art performance across all four benchmark datasets with mean boundary errors below 15 ms. Adaptation and cross-language remapping are effective for languages outside MFA's training distribution, and pronunciation probability modeling and phonological rules provide gains in specific conditions.

语音对齐多语言模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。