构建多语言心理评估数据集,测试大模型跨语言表现
Building Multilingual Datasets for Predicting Mental Health Severity through LLMs: Prospects and Challenges
- 将主流心理数据集翻译成六种语言,支持跨语言评估
- 不同语言下模型表现差异显著,最高误差达35%
- 揭示大模型在医疗场景中的误诊风险,适合临床研究者参考
大型语言模型(LLMs)正被越来越多地应用于心理健康支持系统。然而,关于非英语环境下LLMs有效性研究仍显不足。为此,我们首次将广泛使用的心理评估数据集翻译为希腊语、土耳其语、法语、葡萄牙语、德语和芬兰语六种语言,构建了多语言心理评估数据集,可全面评估模型在多语言环境中检测心理疾病及评估严重程度的能力。通过在GPT与Llama上进行实验,发现即使使用相同翻译数据,各语言间性能存在显著差异。错误分析进一步表明,过度依赖大模型可能导致误诊等医疗风险。该方法大幅降低多语言任务成本,有利于大规模部署。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being integrated into various medical fields, including mental health support systems. However, there is a gap in research regarding the effectiveness of LLMs in non-English mental health support applications. To address this problem, we present a novel multilingual adaptation of widely-used mental health datasets, translated from English into six languages (e.g., Greek, Turkish, French, Portuguese, German, and Finnish). This dataset enables a comprehensive evaluation of LLM performance in detecting mental health conditions and assessing their severity across multiple languages. By experimenting with GPT and Llama, we observe considerable variability in performance across languages, despite being evaluated on the same translated dataset. This inconsistency underscores the complexities inherent in multilingual mental health support, where language-specific nuances and mental health data coverage can affect the accuracy of the models. Through comprehensive error analysis, we emphasize the risks of relying exclusively on LLMs in medical settings (e.g., their potential to contribute to misdiagnoses). Moreover, our proposed approach offers significant cost savings for multilingual tasks, presenting a major advantage for broad-scale implementation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。