arXiv:2508.05102eess.AScs.AI2025-08中稿 · Interspeech 2025被引 3

发现语音克隆模型对构音障碍语音存在识别度优先偏差

Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS

  • 用F5-TTS在TORGO数据集上克隆构音障碍语音
  • 模型更关注可懂度,忽视说话人特征和语调保留
  • 适合关心语音合成公平性的研究者与开发者

构音障碍语音在辅助技术开发中面临数据稀缺挑战。近年来,基于神经网络的零样本语音克隆技术(如F5-TTS)可用于数据增强,但可能引入对构音障碍语音的固有偏见。本文在TORGO数据集上评估F5-TTS在构音障碍语音克隆中的表现,重点分析可懂度、说话人相似度及语调保持能力,并通过不公平度量(如差异影响、差异平等性)评估不同严重程度群体间的差异。结果表明,F5-TTS在构音障碍语音合成中表现出显著偏向于可懂度,而牺牲说话人特征和语调保留。该研究为构建更公平的构音障碍语音合成系统提供了关键洞见。

原文摘要 · Abstract (English)

Dysarthric speech poses significant challenges in developing assistive technologies, primarily due to the limited availability of data. Recent advances in neural speech synthesis, especially zero-shot voice cloning, facilitate synthetic speech generation for data augmentation; however, they may introduce biases towards dysarthric speech. In this paper, we investigate the effectiveness of state-of-the-art F5-TTS in cloning dysarthric speech using TORGO dataset, focusing on intelligibility, speaker similarity, and prosody preservation. We also analyze potential biases using fairness metrics like Disparate Impact and Parity Difference to assess disparities across dysarthric severity levels. Results show that F5-TTS exhibits a strong bias toward speech intelligibility over speaker and prosody preservation in dysarthric speech synthesis. Insights from this study can help integrate fairness-aware dysarthric speech synthesis, fostering the advancement of more inclusive speech technologies.

语音合成公平性克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。