arXiv:2412.16874cs.AIeess.AS2024-12被引 6

融合语音与文本信息,提升口吃症检测与严重程度评估准确率

A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information

  • 用跨注意力机制对齐语音与文本表征,捕捉发音偏差
  • 语音检测准确率达99.53%,严重程度评估达98.12%
  • 适合临床辅助诊断与个性化康复方案设计

自动检测与评估口吃症对精准治疗至关重要。现有研究多聚焦语音模态,本文提出一种融合语音与文本的新方法。通过跨注意力机制学习语音与文本表征间的声学与语言相似性,针对性分析不同严重程度下的发音偏差,显著提升检测与评估精度。所有实验基于UA-Speech口吃语音数据库进行,在说话人依赖与独立、未见词与已见词设置下,检测准确率分别达到99.53%和93.20%,严重程度评估准确率为98.12%和51.97%。结果表明,引入文本提供的语言参考知识,构建了更鲁棒的口吃症检测与评估框架,有望实现更有效的临床诊断。

原文摘要 · Abstract (English)

Automatic detection and severity assessment of dysarthria are crucial for delivering targeted therapeutic interventions to patients. While most existing research focuses primarily on speech modality, this study introduces a novel approach that leverages both speech and text modalities. By employing cross-attention mechanism, our method learns the acoustic and linguistic similarities between speech and text representations. This approach assesses specifically the pronunciation deviations across different severity levels, thereby enhancing the accuracy of dysarthric detection and severity assessment. All the experiments have been performed using UA-Speech dysarthric database. Improved accuracies of 99.53% and 93.20% in detection, and 98.12% and 51.97% for severity assessment have been achieved when speaker-dependent and speaker-independent, unseen and seen words settings are used. These findings suggest that by integrating text information, which provides a reference linguistic knowledge, a more robust framework has been developed for dysarthric detection and assessment, thereby potentially leading to more effective diagnoses.

口吃检测多模态语音分析临床辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。