首个日语全双工对话模型,支持实时语音重叠与回应。
Towards a Japanese Full-duplex Spoken Dialogue System
- 基于英文模型Moshi,分两阶段训练日语对话系统。
- 结合合成数据,自然度和意义性均优于基线模型。
- 适合研究日语实时对话或语音交互的开发者。
全双工语音对话系统能模拟人类对话中的语音重叠和反馈行为,近年来受到广泛关注。然而,针对日语的全双工对话系统研究仍十分有限。本文首次提出公开可用的日语全双工对话模型,基于英文模型Moshi构建。该模型采用两阶段训练:首先在大规模日语语音对话数据上预训练,再在高质量立体语音对话数据上微调。通过引入多流文本转语音生成的合成对话数据,进一步提升性能。实验结果表明,该模型在自然度与语义性方面均优于现有日语基线模型。
原文摘要 · Abstract (English)
Full-duplex spoken dialogue systems, which can model simultaneous bidirectional features of human conversations such as speech overlaps and backchannels, have attracted significant attention recently. However, the study of full-duplex spoken dialogue systems for the Japanese language has been limited, and the research on their development in Japanese remains scarce. In this paper, we present the first publicly available full-duplex spoken dialogue model in Japanese, which is built upon Moshi, a full-duplex dialogue model in English. Our model is trained through a two-stage process: pre-training on a large-scale spoken dialogue data in Japanese, followed by fine-tuning on high-quality stereo spoken dialogue data. We further enhance the model's performance by incorporating synthetic dialogue data generated by a multi-stream text-to-speech system. Evaluation experiments demonstrate that the trained model outperforms Japanese baseline models in both naturalness and meaningfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。