用认知测试分数校准提示,提升语音诊断阿尔茨海默病的准确率。
MMSE-Calibrated Few-Shot Prompting for Alzheimer's Detection
- 用MMSE分数映射提示示例,实现可解释的概率输出。
- 最高达0.86 AUC,超越现有提示方法性能。
- 适合医疗AI研究者和临床辅助系统开发者。
利用大型语言模型进行无训练的阿尔茨海默病检测,基于ADReSS数据集,重新审视零样本提示并研究少样本提示。采用类别平衡协议,通过嵌套交错与严格模式,每类最多使用20个示例。提出两种变体:(i) MMSE-Proxy Prompting:每个少样本示例关联基于MMSE量表的确定性概率,支持计算AUC,达到0.82准确率与0.86 AUC;(ii) Reasoning-augmented Prompting:利用多模态大模型(GPT-5)输入图像、语音转录和MMSE,生成推理过程与对齐概率,评估仅使用转录文本,达0.82准确率与0.83 AUC。据我们所知,这是首个将提示生成概率锚定于MMSE量表,并利用多模态构建增强可解释性的研究。
原文摘要 · Abstract (English)
Prompting large language models is a training-free method for detecting Alzheimer's disease from speech transcripts. Using the ADReSS dataset, we revisit zero-shot prompting and study few-shot prompting with a class-balanced protocol using nested interleave and a strict schema, sweeping up to 20 examples per class. We evaluate two variants achieving state-of-the-art prompting results. (i) MMSE-Proxy Prompting: each few-shot example carries a probability anchored to Mini-Mental State Examination bands via a deterministic mapping, enabling AUC computing; this reaches 0.82 accuracy and 0.86 AUC (ii) Reasoning-augmented Prompting: few-shot examples pool is generated with a multimodal LLM (GPT-5) that takes as input the Cookie Theft image, transcript, and MMSE to output a reasoning and MMSE-aligned probability; evaluation remains transcript-only and reaches 0.82 accuracy and 0.83 AUC. To our knowledge, this is the first ADReSS study to anchor elicited probabilities to MMSE and to use multimodal construction to improve interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。