用合成数据提升多语言认知衰退检测,让模型会解释、跨语言可用。
Cross-Linguistic Persona-Driven Data Synthesis for Robust Multimodal Cognitive Decline Detection
- 通过虚拟患者生成多语言语音数据,缓解临床数据不足问题。
- 在两个基准上实现80.67%和78.46%的宏观F1,优于现有模型。
- 引入思维链微调,让诊断过程可解释,适合临床部署与多语言场景。
基于语音的数字生物标志物为早期识别轻度认知障碍(MCI)提供了可扩展、无创的路径。然而,稳健诊断模型的发展受限于临床数据稀缺及缺乏可解释推理。现有方法常面临跨语言泛化能力差、难以提供透明决策依据的问题。为此,我们提出SynCog框架,结合可控零样本多模态数据合成与思维链(CoT)推理微调。该框架模拟具有不同认知特征的虚拟受试者,有效缓解数据稀缺问题。生成式范式可实现低资源语言下临床语料的快速零样本扩展,显著提升多模态大模型(MLLM)的诊断性能。基于合成数据,采用CoT策略微调基础多模态模型,使模型能显式表达诊断逻辑而非依赖黑箱预测。在ADReSS和ADReSSo基准上的实验表明,用合成表型扩充有限临床数据可达到80.67%和78.46%的宏平均F1,优于当前基线模型。此外,在独立真实中文队列(CIR-E)上评估显示跨语言泛化能力良好,取得48.71%的宏平均F1。这些成果推动了可信赖、跨语言包容的认知评估工具向全球医疗应用迈进。
原文摘要 · Abstract (English)
Speech-based digital biomarkers represent a scalable, non-invasive frontier for the early identification of Mild Cognitive Impairment (MCI). However, the development of robust diagnostic models remains impeded by acute clinical data scarcity and a lack of interpretable reasoning. Current solutions frequently struggle with cross-lingual generalization and fail to provide the transparent rationales essential for clinical trust. To address these barriers, we introduce SynCog, a novel framework integrating controllable zero-shot multimodal data synthesis with Chain-of-Thought (CoT) deduction fine-tuning. Specifically, SynCog simulates diverse virtual subjects with varying cognitive profiles to effectively alleviate clinical data scarcity. This generative paradigm enables the rapid, zero-shot expansion of clinical corpora across diverse languages, effectively bypassing data bottlenecks in low-resource settings and bolstering the diagnostic performance of Multimodal Large Language Models (MLLMs). Leveraging this synthesized dataset, we fine-tune a foundational multimodal backbone using a CoT deduction strategy, empowering the model to explicitly articulate diagnostic thought processes rather than relying on black-box predictions. Extensive experiments on the ADReSS and ADReSSo benchmarks demonstrate that augmenting limited clinical data with synthetic phenotypes yields competitive diagnostic performance, achieving Macro-F1 scores of 80.67% and 78.46%, respectively, outperforming current baseline models. Furthermore, evaluation on an independent real-world Mandarin cohort (CIR-E) demonstrates robust cross-linguistic generalization, attaining a Macro-F1 of 48.71%. These findings constitute a critical step toward providing clinically trustworthy and linguistically inclusive cognitive assessment tools for global healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。