arXiv:2508.04512eess.AS2025-08中稿 · INTERSPEECH 2025被引 1

分析语音痴呆评估的陷阱,发现严重患者结果易被高估

Pitfalls and Limits in Automatic Dementia Assessment

  • 用语音测试评估痴呆,但词命名依赖导致严重患者结果偏高
  • 健康及轻度患者与人工评分相关性明显下降,存在系统偏差
  • 适合关注医疗AI公平性的研究者,提醒警惕数据偏差

当前基于语音的痴呆评估研究主要聚焦于特征提取以预测评估量表,或自动化现有测试流程。多数研究盲目使用公开数据,极少进行细致错误分析,仅关注数值性能。本文对标准化痴呆评估工具Syndrom-Kurz-Test进行深入分析,发现尽管整体与人工评分有高相关性,但因特定数据偏差,严重受损个体的相关性偏高,而健康或轻度受损个体则不然。认知衰退导致语音产出减少,使依赖词命名的评分方式产生过度乐观的结果。此外,测试设计中的备用处理机制会引入额外偏差,偏向特定人群。这些陷阱与数据集中的群体分布无关,需对目标群体进行差异化分析。

原文摘要 · Abstract (English)

Current work on speech-based dementia assessment focuses on either feature extraction to predict assessment scales, or on the automation of existing test procedures. Most research uses public data unquestioningly and rarely performs a detailed error analysis, focusing primarily on numerical performance. We perform an in-depth analysis of an automated standardized dementia assessment, the Syndrom-Kurz-Test. We find that while there is a high overall correlation with human annotators, due to certain artifacts, we observe high correlations for the severely impaired individuals, which is less true for the healthy or mildly impaired ones. Speech production decreases with cognitive decline, leading to overoptimistic correlations when test scoring relies on word naming. Depending on the test design, fallback handling introduces further biases that favor certain groups. These pitfalls remain independent of group distributions in datasets and require differentiated analysis of target groups.

痴呆评估语音分析医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。