评测8个大模型在阿拉伯语心理健康任务中的表现,发现提示工程和少样本学习显著提升效果。
A Comprehensive Evaluation of Large Language Models on Mental Illnesses in Arabic Context
- 设计结构化提示,使多分类任务平均准确率提升14.5%。
- Phi-3.5 MoE在二分类中平衡准确率最高,GPT-4o Mini少样本下多分类准确率提高1.58倍。
- 适合研究阿拉伯语心理健康诊断的大模型应用者参考。
精神健康障碍在阿拉伯世界日益成为重大公共卫生问题,亟需可及的诊断与干预工具。大语言模型(LLMs)虽具潜力,但在阿拉伯语语境下面临标注数据少、语言复杂及翻译偏差等挑战。本研究系统评估了8个大模型,涵盖通用多语言与双语模型,在AraDepSu、Dreaddit、MedMCQA等多样化心理疾病数据集上,探究提示设计、语言配置(原生阿拉伯语/英译阿拉伯语/反向)及少样本提示对诊断性能的影响。结果表明,提示工程显著影响模型表现,主要因指令遵循能力下降;结构化提示在多分类任务中平均优于非结构化提示14.5%。语言配置对性能影响较小,但模型选择至关重要:Phi-3.5 MoE在平衡准确率上表现最优,尤其适用于二分类;Mistral NeMo在严重程度预测任务中均方误差最低。少样本提示持续提升性能,其中GPT-4o Mini在多分类任务中平均准确率提升1.58倍。研究强调提示优化、多语言分析与少样本学习对开发文化敏感且高效的阿拉伯语心理健康大模型工具的重要性。
原文摘要 · Abstract (English)
Mental health disorders pose a growing public health concern in the Arab world, emphasizing the need for accessible diagnostic and intervention tools. Large language models (LLMs) offer a promising approach, but their application in Arabic contexts faces challenges including limited labeled datasets, linguistic complexity, and translation biases. This study comprehensively evaluates 8 LLMs, including general multi-lingual models, as well as bi-lingual ones, on diverse mental health datasets (such as AraDepSu, Dreaddit, MedMCQA), investigating the impact of prompt design, language configuration (native Arabic vs. translated English, and vice versa), and few-shot prompting on diagnostic performance. We find that prompt engineering significantly influences LLM scores mainly due to reduced instruction following, with our structured prompt outperforming a less structured variant on multi-class datasets, with an average difference of 14.5\%. While language influence on performance was modest, model selection proved crucial: Phi-3.5 MoE excelled in balanced accuracy, particularly for binary classification, while Mistral NeMo showed superior performance in mean absolute error for severity prediction tasks. Few-shot prompting consistently improved performance, with particularly substantial gains observed for GPT-4o Mini on multi-class classification, boosting accuracy by an average factor of 1.58. These findings underscore the importance of prompt optimization, multilingual analysis, and few-shot learning for developing culturally sensitive and effective LLM-based mental health tools for Arabic-speaking populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。