arXiv:2602.14584eess.AScs.SD2026-02

用音文对齐技术提升中风后失语症患者的单词命名识别准确率

CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia

  • 将语音与文本映射到统一空间,通过音文匹配识别目标词
  • 在法语患者数据集上最高达90%准确率,显著优于传统方法
  • 适合需要精准评估失语症患者语言能力的研究与临床场景

传统的自动单词命名识别系统难以准确识别中风后失语症患者的词汇,因其存在言语不流畅和发音错误,限制了该群体的可靠自动化评估。本文提出一种基于对比语言-音频预训练(CLAP)的方法,利用音文对齐解决此问题。该方法将单词命名识别视为音频-文本匹配任务,将语音信号与文本提示投影至共享嵌入空间,从而在困难录音条件下仍能识别出患者意图表达的词语。在两个法语中风后失语症患者语音数据集上评估,本方法最高达到90%的识别准确率,优于现有的分类基和自动语音识别基基线。

原文摘要 · Abstract (English)

Conventional automatic word-naming recognition systems struggle to recognize words from post-stroke patients with aphasia because of disfluencies and mispronunciations, limiting reliable automated assessment in this population. In this paper, we propose a Contrastive Language-Audio Pretraining (CLAP) based approach for automatic word-naming recognition to address this challenge by leveraging text-audio alignment. Our approach treats word-naming recognition as an audio-text matching problem, projecting speech signals and textual prompts into a shared embedding space to identify intended words even in challenging recordings. Evaluated on two speech datasets of French post-stroke patients with aphasia, our approach achieves up to 90% accuracy, outperforming existing classification-based and automatic speech recognition-based baselines.

失语症音文对齐CLAP语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。