arXiv:2606.19996cs.SDcs.CL2026-06

用分段语音表征提升中文认知障碍检测精度

Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning

  • 分段处理语音并结合对比学习增强特征表达
  • 在四个中文语料库上三分类任务表现显著提升
  • 适合资源有限的临床筛查场景使用

语音作为一种低成本、无创的数字生物标志物,在认知障碍检测中具有巨大潜力。然而,标注数据有限和跨数据集差异仍是构建稳健语音筛查系统的主要挑战。本文提出一种分段级语音表征学习框架:将语音录音切分为短片段并转为频谱图表示;通过离线与在线数据增强,结合自编码器与对比学习目标,提升在数据稀缺条件下的判别性表征能力。在四个独立的普通话语音数据集上进行实验,二分类与三分类任务均表现出稳定且有竞争力的性能,尤其在临床难度较高的三分类设置中提升明显。消融实验进一步验证了该框架的有效性。结果表明,分段级语音表征学习可为资源受限的临床环境提供可扩展、实用的认知障碍筛查方案。

原文摘要 · Abstract (English)

\noindent\textbf{Background and Objective:} Speech has emerged as a low-cost and non-invasive digital biomarker with considerable potential for cognitive impairment detection. However, limited labeled data and cross-dataset variability remain major challenges for robust speech-based screening systems. \par\noindent\textbf{Methods:} We developed a segment-level representation learning framework for speech-based cognitive impairment detection. Speech recordings were divided into short segments and converted into spectrogram representations. To improve robustness under limited-data conditions, offline and online augmentation strategies were combined with autoencoder-based representation learning and contrastive objectives to enhance discriminative latent representations. \par\noindent\textbf{Results:} Experiments conducted on four independent Mandarin Chinese speech datasets demonstrated stable and competitive performance in both binary and three-class classification tasks, with particularly notable improvements in the clinically challenging three-class setting. Ablation studies further supported the effectiveness of the proposed framework. \par\noindent\textbf{Conclusions:} The findings suggest that segment-level speech representation learning may provide a scalable and practical approach for cognitive impairment screening in resource-constrained clinical settings.

语音识别认知障碍自编码器对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。