arXiv:2606.10675cs.CLeess.AS2026-06

用自监督表示与学习型动态规划实现多语言精准词级对齐

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

  • 融合MMS与自监督音素边界检测器的双表示编码器
  • 在TIMIT和Buckeye上优于MFA和MMS基线,跨语言性能稳定
  • 可直接扩展至MMS支持的1100+语言,无需额外训练

本文提出一种多语言词级强制对齐方法,由对齐编码器和学习型对齐解码器构成。编码器融合来自大规模多语言语音(MMS)模型和自监督音素边界检测器(UnSupSeg)的两种表示,学习在长时序上下文中融合特征并估计词边界概率。解码器为学习型动态规划,结合编码器输出与段落级特征,从MMS和UnSupSeg表示中推断最终词边界。在TIMIT和Buckeye数据集上迭代训练后,该方法在两个数据集上均优于蒙特利尔强制对齐器(MFA)和基于MMS的对齐方法。在未见语言(荷兰语、德语、希伯来语)上表现一致优于或媲美现有方法,表明其具备在不进一步训练的前提下扩展至MMS支持的1100+语言的潜力。

原文摘要 · Abstract (English)

We present a method for accurate multilingual word-level forced alignment, consisting of an alignment encoder and a learned alignment decoder. The encoder integrates two representations: one from the Massively Multilingual Speech (MMS) model and another from a self-supervised phoneme boundary detector (UnSupSeg). It learns to fuse them and to estimate word-boundary probabilities over long temporal contexts. The alignment decoder is a learned dynamic programming that combines encoder outputs with segmental features over the MMS and UnSupSeg representations to infer final word boundaries. Trained iteratively on TIMIT and Buckeye, the proposed approach outperforms Montreal Forced Aligner (MFA) and MMS-based alignment on both datasets. On unseen languages (Dutch, German, and Hebrew), the proposed model achieves performance consistently better than or on par with existing alignment approaches, indicating its potential to scale to 1100+ languages supported by MMS without further training.

语音对齐多语言自监督动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。