首个可实时处理的口音转换模型,保持原音色与语调。
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
- 采用Emformer编码器与优化推理机制实现流式处理
- 在保留说话人特征和韵律的同时提升发音准确性
- 适合实时语音翻译、口语教学等场景
我们提出首个可流式处理的口音转换(AC)模型,能将非母语语音转换为接近母语口音,同时保持说话人身份、语调并改善发音。通过在原有AC架构中引入Emformer编码器和优化推理机制,实现了流式处理。此外,整合了母语文本到语音(TTS)模型以生成理想真值数据,提升训练效率。该流式AC模型性能媲美顶级非流式模型,且延迟稳定,是首个支持实时处理的口音转换系统。
原文摘要 · Abstract (English)
We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronunciation. Our approach enables stream processing by modifying a previous AC architecture with an Emformer encoder and an optimized inference mechanism. Additionally, we integrate a native text-to-speech (TTS) model to generate ideal ground-truth data for efficient training. Our streaming AC model achieves comparable performance to the top AC models while maintaining stable latency, making it the first AC system capable of streaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。