BFA实现毫秒级多语言语音对齐,支持精确边界检测与实时处理。
BFA: Real-time Multilingual Text-to-speech Forced Alignment
- 用无上下文音素编码器+CTC解码,显式建模音素间隙与静音
- 在TIMIT和Buckeye数据集上召回率媲美MFA,且可预测起止边界
- 比MFA快240倍,适合交互式语音应用
我们提出伯恩茅斯强制对齐系统(BFA),结合无上下文通用音素编码器(CUPE)与基于连接时序分类(CTC)的解码器。BFA显式建模音素间间隙与静音,并采用分层解码策略,实现细粒度边界预测。在TIMIT和Buckeye语料库上的评估显示,BFA在宽松容忍度下达到与蒙特利尔强制对齐器(MFA)相当的召回率,同时可预测音素的起始与终止边界,提供更丰富的时序结构。BFA处理速度比MFA快240倍,支持超实时对齐,为以往受限于慢速对齐器的交互式语音应用开辟新可能。
原文摘要 · Abstract (English)
We present Bournemouth Forced Aligner (BFA), a system that combines a Contextless Universal Phoneme Encoder (CUPE) with a connectionist temporal classification (CTC)based decoder. BFA introduces explicit modelling of inter-phoneme gaps and silences and hierarchical decoding strategies, enabling fine-grained boundary prediction. Evaluations on TIMIT and Buckeye corpora show that BFA achieves competitive recall relative to Montreal Forced Aligner at relaxed tolerance levels, while predicting both onset and offset boundaries for richer temporal structure. BFA processes speech up to 240x faster than MFA, enabling faster than real-time alignment. This combination of speed and silence-aware alignment opens opportunities for interactive speech applications previously constrained by slow aligners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。