arXiv:2506.12073eess.AScs.AI2025-06中稿 · Interspeech2025被引 9

用神经网络提升失语症语音与文本对齐精度

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis

  • 基于音素级建模,解决部分对齐与上下文相似性问题
  • 在模拟数据和真实患者数据上均超越现有最佳模型
  • 适合语音病理分析、自动诊断系统开发人员使用

准确对齐失语症语音与其目标文本,对自动化神经退行性语言障碍诊断至关重要。传统方法难以有效建模音素相似性,限制了性能。本文提出 Neural LCS,一种新型的失语症文本-文本及语音-文本对齐方法。该方法通过鲁棒的音素级建模,解决部分对齐与上下文感知相似性映射等关键挑战。我们在大规模模拟数据集(采用先进数据生成技术构建)和真实原发性进行性失语症(PPA)数据上评估该方法,结果表明 Neural LCS 在对齐准确率和失语症语音分割方面显著优于当前最优模型。实验验证了 Neural LCS 在提升失语症自动诊断与分析系统中的潜力,提供了一种更准确且语言学基础坚实的失语症语音对齐方案。

原文摘要 · Abstract (English)

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance. In this work, we propose Neural LCS, a novel approach for dysfluent text-text and speech-text alignment. Neural LCS addresses key challenges, including partial alignment and context-aware similarity mapping, by leveraging robust phoneme-level modeling. We evaluate our method on a large-scale simulated dataset, generated using advanced data simulation techniques, and real PPA data. Neural LCS significantly outperforms state-of-the-art models in both alignment accuracy and dysfluent speech segmentation. Our results demonstrate the potential of Neural LCS to enhance automated systems for diagnosing and analyzing speech disorders, offering a more accurate and linguistically grounded solution for dysfluent speech alignment.

语音对齐失语症分析音素建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。