arXiv:2607.22304cs.LGcs.SD2026-07中稿 · Interspeech 2026

用语音克隆增强临床语音数据,提升低资源语言的情感识别效果

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

论文配图:Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning
图 1 · 摘自论文原文
  • 用8种语音克隆模型生成合成语音,保留关键情感特征
  • 克隆数据训练使日语抑郁焦虑检测准确率超越原始跨语言迁移
  • 适合临床语音数据稀缺场景,尤其对低资源语言有应用价值

语音合成数据增强在语音识别等语言任务中常见,但在需要情感信息的副语言任务中研究较少,尤其是标注成本高、患者群体代表性不足的临床场景。语音克隆是一种潜在的增强方法,但以往评估多关注语音可懂度(WER)或说话人相似性(SS),未考察其是否保留下游任务所需的副语言信号。本研究在五个公共与临床数据集上对八种语音克隆模型进行基准测试,发现多数模型仅造成轻微信号损失。进一步将英语临床语音克隆为日语后,基于克隆数据训练的模型在真实日语语音上进行抑郁与焦虑检测时表现优于直接跨语言迁移,表明语音克隆是低资源语言临床语音数据增强的可行路径。

原文摘要 · Abstract (English)

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear whether these preserve the paralinguistic signal such tasks depend on. We benchmark eight voice cloning models on five paralinguistic tasks across public and clinical datasets, showing most preserve signal with modest degradation. We then clone English clinical speech into Japanese and find that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.

语音克隆临床语音情感识别数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。