arXiv:2503.23873eess.AScs.SD2025-03被引 1

用ChatGPT-4o实现少样本语音病理检测,兼顾准确与可解释性。

Exploring In-Context Learning Capabilities of ChatGPT for Pathological Speech Detection

  • 基于少样本上下文学习,利用多模态大模型分析语音数据。
  • 在少样本设置下达到高检测准确率,且能生成决策解释。
  • 适合临床场景中需要透明诊断依据的研究与应用。

自动语音病理检测方法已展现出良好效果,有望作为昂贵传统方法的替代诊断工具。尽管这些方法准确性高,但缺乏可解释性,限制了其在临床中的应用。本文研究了多模态大语言模型(特别是ChatGPT-4o)在少样本上下文学习设置下进行自动语音病理检测的可行性。实验结果表明,该方法不仅性能出色,还能提供决策解释,提升模型可解释性。为进一步理解其有效性,我们进行了消融实验,分析输入类型和系统提示对结果的影响。研究结果凸显了多模态大模型在自动语音病理检测领域的潜力。

原文摘要 · Abstract (English)

Automatic pathological speech detection approaches have shown promising results, gaining attention as potential diagnostic tools alongside costly traditional methods. While these approaches can achieve high accuracy, their lack of interpretability limits their applicability in clinical practice. In this paper, we investigate the use of multimodal Large Language Models (LLMs), specifically ChatGPT-4o, for automatic pathological speech detection in a few-shot in-context learning setting. Experimental results show that this approach not only delivers promising performance but also provides explanations for its decisions, enhancing model interpretability. To further understand its effectiveness, we conduct an ablation study to analyze the impact of different factors, such as input type and system prompts, on the final results. Our findings highlight the potential of multimodal LLMs for further exploration and advancement in automatic pathological speech detection.

语音病理大模型可解释性少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。