arXiv:2602.02440cs.CL2026-02被引 5

评测大模型在多语言心理健康任务中的表现,发现翻译质量显著影响效果

Large Language Models for Mental Health: A Multilingual Evaluation

  • 在八种语言的八组心理健康数据上对比了开源与闭源大模型的表现
  • 微调后的开源模型在多个数据集上达到或超过现有最佳水平
  • 机器翻译数据上的性能下降明显,尤其受语言类型差异影响

大型语言模型(LLMs)在自然语言处理任务中表现出色,但在多语言语境下,尤其是在心理健康领域,其表现尚未得到充分研究。本文评估了专有和开源的LLMs在八种不同语言的心理健康数据集及其机器翻译(MT)版本上的表现,并与不使用LLM的传统NLP基线方法进行了对比。测试涵盖零样本、少样本和微调三种设置。此外,我们评估了不同语言家族和语言类型下的翻译质量,以分析其对LLM性能的影响。结果显示,专有LLMs和微调后的开源LLMs在多个数据集上取得了具有竞争力的F1分数,常常超越当前最优结果。然而,机器翻译数据上的性能普遍较低,且下降程度因语言和语言类型而异。这一差异凸显了大模型在非英语语言中处理心理健康任务的优势,也揭示了当翻译质量导致结构或词汇错配时,其局限性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have remarkable capabilities across NLP tasks. However, their performance in multilingual contexts, especially within the mental health domain, has not been thoroughly explored. In this paper, we evaluate proprietary and open-source LLMs on eight mental health datasets in various languages, as well as their machine-translated (MT) counterparts. We compare LLM performance in zero-shot, few-shot, and fine-tuned settings against conventional NLP baselines that do not employ LLMs. In addition, we assess translation quality across language families and typologies to understand its influence on LLM performance. Proprietary LLMs and fine-tuned open-source LLMs achieve competitive F1 scores on several datasets, often surpassing state-of-the-art results. However, performance on MT data is generally lower, and the extent of this decline varies by language and typology. This variation highlights both the strengths of LLMs in handling mental health tasks in languages other than English and their limitations when translation quality introduces structural or lexical mismatches.

大模型心理健康多语言翻译质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。