利用语音停顿同步实现非洲语言无字幕语音翻译,提升准确率并降低错误。
Phonology-Guided Speech-to-Speech Translation for African Languages
- 基于跨语言停顿同步设计动态规划对齐算法,无需转录文本
- 新模型在多语言语音翻译中将误译率降低至5.3%,实时性达RTF=1.02
- 适用于低资源语言,提供无需转录的评测体系,适合语音翻译研究者
我们提出一种以语调为引导的语音到语音翻译(S2ST)框架,通过跨语言停顿同步,在无需转录的情况下实现语音对齐与翻译。分析涵盖五种语言、6000小时的东非新闻语料库发现,同语系语言对之间的停顿方差低30%–40%,起始/结束时间相关性超过跨语系对的3倍以上。据此提出SPaDA动态规划对齐算法,融合静默一致性、速率同步与语义相似性,使对齐F1提升3–4点,减少最多38%的虚假匹配。基于SPaDA对齐片段,训练了受冻结语义与说话人编码器外部梯度引导的扩散模型SegUniDiff。该模型在CVSS-C上达到30.3的BLEU分,优于单位基线(28.9),说话人错误率从12.5%降至5.3%,实时因子(RTF)为1.02。为支持低资源场景评估,我们还发布三层次、无转录的BLEU评测套件(M1–M3),与人工评价高度相关。结果表明,多语言语音中的韵律线索可作为可扩展、非自回归语音翻译的可靠基础。
原文摘要 · Abstract (English)
We present a prosody-guided framework for speech-to-speech translation (S2ST) that aligns and translates speech \emph{without} transcripts by leveraging cross-linguistic pause synchrony. Analyzing a 6{,}000-hour East African news corpus spanning five languages, we show that \emph{within-phylum} language pairs exhibit 30--40\% lower pause variance and over 3$\times$ higher onset/offset correlation compared to cross-phylum pairs. These findings motivate \textbf{SPaDA}, a dynamic-programming alignment algorithm that integrates silence consistency, rate synchrony, and semantic similarity. SPaDA improves alignment $F_1$ by +3--4 points and eliminates up to 38\% of spurious matches relative to greedy VAD baselines. Using SPaDA-aligned segments, we train \textbf{SegUniDiff}, a diffusion-based S2ST model guided by \emph{external gradients} from frozen semantic and speaker encoders. SegUniDiff matches an enhanced cascade in BLEU (30.3 on CVSS-C vs.\ 28.9 for UnitY), reduces speaker error rate (EER) from 12.5\% to 5.3\%, and runs at an RTF of 1.02. To support evaluation in low-resource settings, we also release a three-tier, transcript-free BLEU suite (M1--M3) that correlates strongly with human judgments. Together, our results show that prosodic cues in multilingual speech provide a reliable scaffold for scalable, non-autoregressive S2ST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。