arXiv:2508.03360cs.AI2025-08被引 3

首个跨语言语音认知评估大模型基准,测试多语言场景下模型表现

CogBench: A Large Language Model Benchmark for Multilingual Speech-Based Cognitive Impairment Assessment

  • 构建统一多模态流程,评估大模型在英、中文语音数据上的表现
  • 大模型经思维链提示后适应性更强,但效果依赖提示设计
  • 轻量化微调(LoRA)显著提升模型在目标领域的泛化能力

从自发性语音自动评估认知障碍为早期筛查提供了有前景的无创途径。然而,现有方法在跨语言和临床场景部署时普遍缺乏泛化能力,限制了实际应用。本文提出CogBench,首个用于评估大语言模型(LLM)在语音基认知障碍评估中跨语言与跨站点泛化能力的基准。采用统一多模态流程,在涵盖英语和普通话的三个语音数据集——ADReSSo、NCMMSC2021-AD及新采集的CIR-E测试集上评估模型性能。结果显示,传统深度学习模型在跨域迁移时性能大幅下降;而配备思维链提示的大模型表现出更好适应性,但其性能仍对提示设计敏感。此外,通过低秩适配(LoRA)进行轻量级微调可显著提升目标域的泛化性能。这些发现为构建临床可用且语言鲁棒的语音认知评估工具迈出关键一步。

原文摘要 · Abstract (English)

Automatic assessment of cognitive impairment from spontaneous speech offers a promising, non-invasive avenue for early cognitive screening. However, current approaches often lack generalizability when deployed across different languages and clinical settings, limiting their practical utility. In this study, we propose CogBench, the first benchmark designed to evaluate the cross-lingual and cross-site generalizability of large language models (LLMs) for speech-based cognitive impairment assessment. Using a unified multimodal pipeline, we evaluate model performance on three speech datasets spanning English and Mandarin: ADReSSo, NCMMSC2021-AD, and a newly collected test set, CIR-E. Our results show that conventional deep learning models degrade substantially when transferred across domains. In contrast, LLMs equipped with chain-of-thought prompting demonstrate better adaptability, though their performance remains sensitive to prompt design. Furthermore, we explore lightweight fine-tuning of LLMs via Low-Rank Adaptation (LoRA), which significantly improves generalization in target domains. These findings offer a critical step toward building clinically useful and linguistically robust speech-based cognitive assessment tools.

大模型语音评估认知障碍跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。