arXiv:2506.04397eess.AScs.CL2025-06中稿 · Interspeech 2025被引 5

用大模型重建失语者旧嗓音,让沟通更自然。

Can we reconstruct a dysarthric voice with the large speech model Parler TTS?

  • 用Parler TTS模型复现失语前的说话声音
  • 生成语音可听懂但身份一致性不足
  • 适合语音康复与个性化助残研究者

言语障碍会使患者交流困难甚至无法沟通。个性化文本转语音(TTS)可作为有效的辅助工具。本文尝试利用大模型重建失语前的语音特征,使用状态先进的大模型Parler TTS,生成患者发病前的声音近似。我们构建了一个数据集并标注了说话人信息和语音可理解性,用于微调模型。结果表明,模型能够学习该挑战性数据的分布,但对语音可理解性和说话人身份的一致性控制能力有限。本文提出未来改进方向,以增强此类模型在语音重建任务中的可控性。

原文摘要 · Abstract (English)

Speech disorders can make communication hard or even impossible for those who develop them. Personalised Text-to-Speech is an attractive option as a communication aid. We attempt voice reconstruction using a large speech model, with which we generate an approximation of a dysarthric speaker's voice prior to the onset of their condition. In particular, we investigate whether a state-of-the-art large speech model, Parler TTS, can generate intelligible speech while maintaining speaker identity. We curate a dataset and annotate it with relevant speaker and intelligibility information, and use this to fine-tune the model. Our results show that the model can indeed learn to generate from the distribution of this challenging data, but struggles to control intelligibility and to maintain consistent speaker identity. We propose future directions to improve controllability of this class of model, for the voice reconstruction task.

语音合成失语康复大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。