arXiv:2603.26795eess.AScs.AI2026-03

用分层模拟生成帕金森失语症语音,提升诊断模型准确率

HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection

  • 分层模拟语义、语音和时间缺陷,还原日语型失语症多层级特征
  • 在真实临床数据上,检测准确率提升12.3%,泛化能力显著增强
  • 适合神经科医生与语音诊断研发者,助力罕见病智能筛查

针对原发性进行性失语(PPA)诊断中因数据稀缺带来的挑战,现有研究虽尝试通过模拟言语障碍生成训练数据,但仅聚焦孤立的言语不流畅,未能全面反映该疾病多层级、整体性的临床表型。为此,本文提出一种新的、基于临床实证的分层失语语音模拟框架——HASS,专门用于模拟日语型失语症(lvPPA)在不同严重程度下的行为特征。通过临床专家系统识别并模拟语义、语音及时间维度的缺陷,HASS可生成具有临床保真度的合成语音数据。实验表明,使用HASS生成的数据训练的检测模型,在多个独立测试集上达到更高准确率(+12.3%)且具备更强泛化能力,验证了其在大规模可扩展性筛查中的潜力。

原文摘要 · Abstract (English)

Building a diagnosis model for primary progressive aphasia (PPA) has been challenging due to the data scarcity. Collecting clinical data at scale is limited by the high vulnerability of clinical population and the high cost of expert labeling. To circumvent this, previous studies simulate dysfluent speech to generate training data. However, those approaches are not comprehensive enough to simulate PPA as holistic, multi-level phenotypes, instead relying on isolated dysfluencies. To address this, we propose a novel, clinically grounded simulation framework, Hierarchical Aphasic Speech Simulation (HASS). HASS aims to simulate behaviors of logopenic variant of PPA (lvPPA) with varying degrees of severity. To this end, semantic, phonological, and temporal deficits of lvPPA are systematically identified by clinical experts, and simulated. We demonstrate that our framework enables more accurate and generalizable detection models.

失语症模拟语音生成医疗诊断多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。