用大模型分析对话中的语言特征,提升阿尔茨海默病早期诊断准确率。
DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
- 利用大模型提取对话中的细粒度语言线索,过滤噪声信息。
- 通过组合生成新语音文本,解决数据稀缺问题,准确率提升11%。
- 适合做医疗语言分析与智能诊断的研究者参考。
阿尔茨海默病(AD)影响全球5000万人,早期识别关键标志物对及时干预至关重要。语言障碍是认知衰退最早的表现之一,可通过患者与医生的对话检测。然而,对话中常混杂模糊、噪声和无关信息,且真实AD语音样本少、风格差异大,制约了模型鲁棒性。为此,我们提出DECT,一种基于大语言模型(LLM)的细粒度语言分析与标签保持/切换的数据生成方法。首先,利用LLM的总结能力从嘈杂转录文本中提炼关键认知-语言信息;其次,借助LLM固有的语言知识从非结构化音频转录中提取语言标记;再次,运用其组合能力生成包含多样语言模式的模拟语音文本,缓解数据稀缺问题;最后,使用增强后的文本数据微调检测模型。在DementiaBank数据集上,相比基线模型,DECT将诊断准确率提升11%。
原文摘要 · Abstract (English)
Alzheimer's Disease (AD) is an irreversible neurodegenerative disease affecting 50 million people worldwide. Low-cost, accurate identification of key markers of AD is crucial for timely diagnosis and intervention. Language impairment is one of the earliest signs of cognitive decline, which can be used to discriminate AD patients from normal control individuals. Patient-interviewer dialogues may be used to detect such impairments, but they are often mixed with ambiguous, noisy, and irrelevant information, making the AD detection task difficult. Moreover, the limited availability of AD speech samples and variability in their speech styles pose significant challenges in developing robust speech-based AD detection models. To address these challenges, we propose DECT, a novel speech-based domain-specific approach leveraging large language models (LLMs) for fine-grained linguistic analysis and label-switched label-preserved data generation. Our study presents four novelties: We harness the summarizing capabilities of LLMs to identify and distill key Cognitive-Linguistic information from noisy speech transcripts, effectively filtering irrelevant information. We leverage the inherent linguistic knowledge of LLMs to extract linguistic markers from unstructured and heterogeneous audio transcripts. We exploit the compositional ability of LLMs to generate AD speech transcripts consisting of diverse linguistic patterns to overcome the speech data scarcity challenge and enhance the robustness of AD detection models. We use the augmented AD textual speech transcript dataset and a more fine-grained representation of AD textual speech transcript data to fine-tune the AD detection model. The results have shown that DECT demonstrates superior model performance with an 11% improvement in AD detection accuracy on the datasets from DementiaBank compared to the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。