构建多语言语音翻译数据集,让翻译更传情达意。
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
- 基于电影音频构建精准匹配的多语言语音对齐数据集。
- 融合多种韵律迁移技术,显著保留源语音的情感特征。
- 适合关注情感表达与自然语音生成的研究者使用。
当前语音到语音翻译(S2ST)研究主要聚焦于翻译准确性和语音自然度,常忽视语气、情感等副语言信息在交流中的关键作用。为此,本研究从多部电影音频中精心构建了一个多语言数据集,每对语音在副语言信息和时长上均精确匹配。通过整合多种韵律迁移技术,旨在实现高准确率、自然流畅且富含副语言细节的翻译效果。实验结果表明,该模型在保持高翻译准确性和语音自然度的同时,显著保留了源语音中的副语言信息。
原文摘要 · Abstract (English)
Current research in speech-to-speech translation (S2ST) primarily concentrates on translation accuracy and speech naturalness, often overlooking key elements like paralinguistic information, which is essential for conveying emotions and attitudes in communication. To address this, our research introduces a novel, carefully curated multilingual dataset from various movie audio tracks. Each dataset pair is precisely matched for paralinguistic information and duration. We enhance this by integrating multiple prosody transfer techniques, aiming for translations that are accurate, natural-sounding, and rich in paralinguistic details. Our experimental results confirm that our model retains more paralinguistic information from the source speech while maintaining high standards of translation accuracy and naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。