针对汉语方言数据稀缺问题,提出端到端对话生成模型提升方言语音自然度。
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

- 构建可扩展方言语音对话合成流程,解决数据不足难题
- 采用两阶段后训练与自对齐语音监督,提升语义一致性
- 适用于低资源方言研究及语音助手等实际应用
当前端到端语音对话模型主要针对主流语言优化,在低资源方言场景下表现受限,主要因方言语音数据稀缺。此外,方言适配过程中,模型的语义表示空间持续演化,而传统语音监督保持不变,导致隐藏表示与语音目标间语义不一致,影响语音稳定性和自然度。为此,本文提出 DialectS2S,一种面向汉语方言的端到端语音对话模型。首先构建可扩展的方言语音对话合成流水线以高效生成数据;进一步提出两阶段后训练策略,结合自对齐语音监督,使语音监督内容与模型演化的语义表示对齐,从而提升方言语音生成质量。实验结果表明,DialectS2S 在多个汉语方言上均显著优于现有基线,大幅改善方言一致性、回复质量和语音可懂度。本工作为低资源方言场景下的端到端语音对话建模提供了高效且可扩展的解决方案。为促进后续研究与应用,我们开源了 DialectS2S 框架,包含模型权重、训练数据集及微调代码。
原文摘要 · Abstract (English)
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。