arXiv:2602.07978cs.CL2026-02

用合成数据提升多语言认知衰退检测,让模型会解释、跨语言可用。

Cross-Linguistic Persona-Driven Data Synthesis for Robust Multimodal Cognitive Decline Detection

  • 通过虚拟患者生成多语言语音数据,缓解临床数据不足问题。
  • 在两个基准上实现80.67%和78.46%的宏观F1,优于现有模型。
  • 引入思维链微调,让诊断过程可解释,适合临床部署与多语言场景。

基于语音的数字生物标志物为早期识别轻度认知障碍(MCI)提供了可扩展、无创的路径。然而,稳健诊断模型的发展受限于临床数据稀缺及缺乏可解释推理。现有方法常面临跨语言泛化能力差、难以提供透明决策依据的问题。为此,我们提出SynCog框架,结合可控零样本多模态数据合成与思维链(CoT)推理微调。该框架模拟具有不同认知特征的虚拟受试者,有效缓解数据稀缺问题。生成式范式可实现低资源语言下临床语料的快速零样本扩展,显著提升多模态大模型(MLLM)的诊断性能。基于合成数据,采用CoT策略微调基础多模态模型,使模型能显式表达诊断逻辑而非依赖黑箱预测。在ADReSS和ADReSSo基准上的实验表明,用合成表型扩充有限临床数据可达到80.67%和78.46%的宏平均F1,优于当前基线模型。此外,在独立真实中文队列(CIR-E)上评估显示跨语言泛化能力良好,取得48.71%的宏平均F1。这些成果推动了可信赖、跨语言包容的认知评估工具向全球医疗应用迈进。

原文摘要 · Abstract (English)

Speech-based digital biomarkers represent a scalable, non-invasive frontier for the early identification of Mild Cognitive Impairment (MCI). However, the development of robust diagnostic models remains impeded by acute clinical data scarcity and a lack of interpretable reasoning. Current solutions frequently struggle with cross-lingual generalization and fail to provide the transparent rationales essential for clinical trust. To address these barriers, we introduce SynCog, a novel framework integrating controllable zero-shot multimodal data synthesis with Chain-of-Thought (CoT) deduction fine-tuning. Specifically, SynCog simulates diverse virtual subjects with varying cognitive profiles to effectively alleviate clinical data scarcity. This generative paradigm enables the rapid, zero-shot expansion of clinical corpora across diverse languages, effectively bypassing data bottlenecks in low-resource settings and bolstering the diagnostic performance of Multimodal Large Language Models (MLLMs). Leveraging this synthesized dataset, we fine-tune a foundational multimodal backbone using a CoT deduction strategy, empowering the model to explicitly articulate diagnostic thought processes rather than relying on black-box predictions. Extensive experiments on the ADReSS and ADReSSo benchmarks demonstrate that augmenting limited clinical data with synthetic phenotypes yields competitive diagnostic performance, achieving Macro-F1 scores of 80.67% and 78.46%, respectively, outperforming current baseline models. Furthermore, evaluation on an independent real-world Mandarin cohort (CIR-E) demonstrates robust cross-linguistic generalization, attaining a Macro-F1 of 48.71%. These findings constitute a critical step toward providing clinically trustworthy and linguistically inclusive cognitive assessment tools for global healthcare.

认知衰退多模态合成数据可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。