无需文字转录,用语音嵌入实现高质量语音对齐。
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
- 基于语音段嵌入单调对齐,不依赖文本转录。
- 在3000小时无标注数据上生成约1000小时高质量对齐数据。
- 比现有方法更鲁棒,且仅用1/8数据达到同等翻译性能。
我们提出Speech Vecalign,一种无需文本转录的并行语音文档对齐方法,通过单调对齐语音段嵌入实现。相比基线方法Global Mining,该方法生成更长的语音到语音对齐;相比Local Mining,噪声更少。将Speech Vecalign应用于VoxPopuli中的3000小时无标注英德双语语音数据,获得约1000小时高质量对齐数据。在此基础上训练的英德语音到语音翻译模型,相较Global Mining在英译德和德译英任务上分别提升0.37和0.18 ASR-BLEU。尽管仅使用原始数据的1/8,模型性能仍可媲美甚至超过SpeechMatrix。
原文摘要 · Abstract (English)
We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. Compared to the baseline method Global Mining, a variant of speech mining, Speech Vecalign produces longer speech-to-speech alignments. It also demonstrates greater robustness than Local Mining, another speech mining variant, as it produces less noise. We applied Speech Vecalign to 3,000 hours of unlabeled parallel English-German (En-De) speech documents from VoxPopuli, yielding about 1,000 hours of high-quality alignments. We then trained En-De speech-to-speech translation models on the aligned data. Speech Vecalign improves the En-to-De and De-to-En performance over Global Mining by 0.37 and 0.18 ASR-BLEU, respectively. Moreover, our models match or outperform SpeechMatrix model performance, despite using 8 times fewer raw speech documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。