arXiv:2605.16077cs.CL2026-05

用大模型生成逼真语音数据,提升认知评分预测准确率

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

论文配图:Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
图 1 · 摘自论文原文
  • 用GPT-5基于书面回答生成多种风格的口语化叙述
  • 相似性引导的数据增强使低分人群误差减少37%
  • 适合临床语音分析中样本少、类别不均衡的研究者

由于数据集规模有限和类别不平衡,从自发语音中准确评估认知衰退仍具挑战。本文提出一种大语言模型(LLM)驱动的数据增强框架,以改善基于语音的认知评分预测。实验基于一个日语语料库,每位参与者提供同一临床提示的自发口头叙述和书面回应。书面回应作为语义锚点,使用GPT-5生成多个不同风格的类口语独白。随后,利用偏最小二乘回归模型,基于Sentence-BERT语音嵌入预测广泛用于日本的认知筛查工具——长谷川痴呆量表(Hasegawa Dementia Scale)得分。研究比较两种增强策略:随机类平衡选择,带来中等但不稳定的提升;相似性引导类平衡选择,优先选取语义相近的合成样本,实现更一致的性能提升,并显著降低少数群体(低分者)的预测误差,同时保持多数群体表现。结果表明,语义引导的LLM增强是解决类别不平衡、提升临床语音分析数据效率的可行方法。

原文摘要 · Abstract (English)

Accurate assessment of cognitive decline from spontaneous speech remains challenging due to limited dataset size and class imbalance. In this work, we propose a large language model (LLM)-driven data augmentation framework to improve the prediction of cognitive scores from speech. Experiments are conducted on a Japanese corpus in which each participant provides both a spontaneous oral narrative and a written response to the same clinical prompt. The written responses serve as semantic anchors to generate multiple oral-like monologues in different styles using GPT-5. We then predict Hasegawa Dementia Scale scores, a widely used cognitive screening tool in Japan, using a Partial Least Squares regression model trained on Sentence-BERT speech embeddings. We investigate two augmentation strategies: random class-balanced selection, which yields moderate but unstable improvements, and similarity-guided class-balanced selection. The latter prioritizes semantically close synthetic samples, leading to more consistent improvements and substantially reducing prediction error for minority low-score participants while maintaining performance for the majority group. Overall, our findings demonstrate the potential of semantically guided LLM-driven augmentation as a principled approach for addressing class imbalance and improving data efficiency in clinical speech analysis.

语音分析大模型认知评估数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。