arXiv:2501.15858cs.CLcs.SD2025-01

用AI构建跨语言失语语音可懂度评估框架,突破英语主导局限。

Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech

  • 先用通用声学模型提取失语特征,再结合目标语言特点判断可懂度。
  • 解决数据少、标注难、语言规律不足等跨语言评估核心难题。
  • 适合语音病理学、多语言AI医疗研究者参考应用。

目的:语音可懂度是失语症评估与管理的关键指标,但现有研究和临床实践多集中于英语,限制了其在其他语言中的适用性。本文提出一个概念框架,并展示如何利用人工智能(AI)推进跨语言失语语音可懂度评估。方法:我们设计了一个双层框架——首先通过通用语音模型将失语语音编码为声学-音位表征,随后由语言特定的可懂度评估模型在目标语言的音系或语调结构中解析这些表征。同时,识别出跨语言评估的主要障碍,包括数据稀缺、标注复杂及对失语语音的语言学理解不足,并提出相应的AI驱动解决方案。结论:提升跨语言失语语音可懂度评估需要兼顾效率与可扩展性的模型,同时受限于语言规则以保证准确性与语言敏感性。近期的AI进展为此类整合提供了基础工具,推动未来向可泛化且具语言学洞察的评估框架发展。

原文摘要 · Abstract (English)

Purpose: Speech intelligibility is a critical outcome in the assessment and management of dysarthria, yet most research and clinical practices have focused on English, limiting their applicability across languages. This commentary introduces a conceptual framework--and a demonstration of how it can be implemented--leveraging artificial intelligence (AI) to advance cross-language intelligibility assessment of dysarthric speech. Method: We propose a two-tiered conceptual framework consisting of a universal speech model that encodes dysarthric speech into acoustic-phonetic representations, followed by a language-specific intelligibility assessment model that interprets these representations within the phonological or prosodic structures of the target language. We further identify barriers to cross-language intelligibility assessment of dysarthric speech, including data scarcity, annotation complexity, and limited linguistic insights into dysarthric speech, and outline potential AI-driven solutions to overcome these challenges. Conclusion: Advancing cross-language intelligibility assessment of dysarthric speech necessitates models that are both efficient and scalable, yet constrained by linguistic rules to ensure accurate and language-sensitive assessment. Recent advances in AI provide the foundational tools to support this integration, shaping future directions toward generalizable and linguistically informed assessment frameworks.

语音病理AI评估跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。