用语音助手长期记录命令,通过大模型迭代优化检测认知衰退
Cog-TiPRO: Iterative Prompt Refinement with LLMs to Detect Cognitive Decline via Longitudinal Voice Assistant Commands
- 用大模型迭代优化提示词,提取语音命令中的语言特征
- 18个月数据中识别出轻度认知障碍准确率达73.80%
- 适合关注老龄化健康监测的医疗与AI研究者
早期发现认知衰退对延缓神经退行性疾病进展至关重要。传统诊断依赖耗时的临床评估,难以实现频繁监测。本初步研究探讨语音助手系统(VAS)作为非侵入式工具,通过纵向分析语音命令中的言语模式来检测认知衰退。在18个月期间,我们收集了35名老年人的语音命令数据,其中15人每日在家中使用语音助手。为应对短句、无结构、噪声大的挑战,提出Cog-TiPRO框架:(1)利用大模型驱动的迭代提示优化进行语言特征提取,(2)基于HuBERT的声学特征提取,(3)基于transformer的时间建模。采用iTransformer,该方法在检测轻度认知障碍(MCI)上达到73.80%准确率和72.67% F1-score,优于基线27.13%。通过大模型方法,我们识别出认知衰退个体日常命令使用的独特语言特征。
原文摘要 · Abstract (English)
Early detection of cognitive decline is crucial for enabling interventions that can slow neurodegenerative disease progression. Traditional diagnostic approaches rely on labor-intensive clinical assessments, which are impractical for frequent monitoring. Our pilot study investigates voice assistant systems (VAS) as non-invasive tools for detecting cognitive decline through longitudinal analysis of speech patterns in voice commands. Over an 18-month period, we collected voice commands from 35 older adults, with 15 participants providing daily at-home VAS interactions. To address the challenges of analyzing these short, unstructured and noisy commands, we propose Cog-TiPRO, a framework that combines (1) LLM-driven iterative prompt refinement for linguistic feature extraction, (2) HuBERT-based acoustic feature extraction, and (3) transformer-based temporal modeling. Using iTransformer, our approach achieves 73.80% accuracy and 72.67% F1-score in detecting MCI, outperforming its baseline by 27.13%. Through our LLM approach, we identify linguistic features that uniquely characterize everyday command usage patterns in individuals experiencing cognitive decline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。