arXiv:2409.15687cs.AI2024-09被引 11

评测33个大模型在精神疾病识别中的表现,发现大模型有潜力但受伦理限制。

A Comprehensive Evaluation of Large Language Models on Mental Illnesses

  • 对比33个大模型在社交数据上的零样本和少样本能力
  • GPT-4等模型二分类准确率达85%,Phi-3-mini误差降1.3点
  • 新模型如Llama 3.1 405b知识测评准确率91.2%,适合临床辅助

大型语言模型(LLMs)在医疗等领域展现出潜力,有望通过可扩展、易获取的解决方案变革心理健康应用。本研究对33个参数量从20亿到4050亿以上的现代大模型,在六大数据集上使用社交媒体数据,评估其在精神健康关键任务中的表现,涵盖二元障碍检测、障碍严重程度评估和精神卫生知识测试。评估覆盖GPT-4、Llama 3、Claude、Gemma、Gemini、Phi-3等模型,考察其零样本(ZS)与少样本(FS)能力。结果显示,GPT-4和Llama 3在二元障碍检测中表现优异,部分数据集准确率达85%;少样本学习显著提升严重程度评估,使Phi-3-mini模型的平均绝对误差(MAE)降低1.3点。最新模型如Llama 3.1 405b在精神卫生知识评估中达到91.2%准确率。提示工程对各任务性能提升至关重要。然而,多数模型提供方的伦理限制阻碍其对敏感问题的回应,影响全面评估。该研究揭示了大模型在精神健康领域的潜力与局限,为未来精神病学应用提供重要参考。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown promise in various domains, including healthcare, with significant potential to transform mental health applications by enabling scalable and accessible solutions. This study aims to provide a comprehensive evaluation of 33 LLMs, ranging from 2 billion to 405+ billion parameters, in performing key mental health tasks using social media data across six datasets. To our knowledge, this represents the largest-scale systematic evaluation of modern LLMs for mental health applications. Models such as GPT-4, Llama 3, Claude, Gemma, Gemini, and Phi-3 were assessed for their zero-shot (ZS) and few-shot (FS) capabilities across three tasks: binary disorder detection, disorder severity evaluation, and psychiatric knowledge assessment. Key findings revealed that models like GPT-4 and Llama 3 exhibited superior performance in binary disorder detection, achieving accuracies up to 85% on certain datasets, while FS learning notably enhanced disorder severity evaluations, reducing the Mean Absolute Error (MAE) by 1.3 points for the Phi-3-mini model. Recent models, such as Llama 3.1 405b, demonstrated exceptional psychiatric knowledge assessment accuracy at 91.2%, while prompt engineering played a crucial role in improving performance across tasks. However, the ethical constraints imposed by many LLM providers limit their ability to respond to sensitive queries, hampering comprehensive performance evaluations. This work highlights both the capabilities and limitations of LLMs in mental health contexts, offering valuable insights for future applications in psychiatry.

大模型精神健康评估少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。