arXiv:2502.15418cs.CL2025-02被引 12

构建心理健康的多选题数据集,评估大模型在四大领域的真实问答能力。

MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models

  • 基于医学文献构建四类问题的问答对,覆盖焦虑、抑郁等核心领域。
  • 产出2475个专家验证的黄金标准数据与约5.6万条伪标签数据。
  • 适合评估大模型在心理健康领域的知识理解与推理能力。

心理健康问题在全球范围内日益严峻,抑郁症、焦虑症等愈发普遍。尽管大语言模型(LLMs)已在医疗问答中广泛应用,但缺乏针对心理健康领域的标准化评测数据集。本文提出一个新型多选题数据集MHQA(Mental Health Question Answering),用于评估语言模型在心理健康方面的表现。现有心理健康数据集多聚焦于分类任务,而MHQA则涵盖焦虑、抑郁、创伤及强迫/强迫性问题四大关键领域,包含事实型、诊断型、预后型和预防型四类问题。数据主要源自PubMed摘要,通过严格的基于大模型的信息识别与筛选流程,结合后验验证标准生成高质量问答对。最终,数据集包含2,475个专家验证的黄金标准样本(MHQA-gold)以及约56.1k条使用外部医学参考生成的伪标签样本。我们报告了不同大模型在该数据集上的F1得分,并进行了少样本学习与微调实验,分析模型表现差异。

原文摘要 · Abstract (English)

Mental health remains a challenging problem all over the world, with issues like depression, anxiety becoming increasingly common. Large Language Models (LLMs) have seen a vast application in healthcare, specifically in answering medical questions. However, there is a lack of standard benchmarking datasets for question answering (QA) in mental health. Our work presents a novel multiple choice dataset, MHQA (Mental Health Question Answering), for benchmarking Language models (LMs). Previous mental health datasets have focused primarily on text classification into specific labels or disorders. MHQA, on the other hand, presents question-answering for mental health focused on four key domains: anxiety, depression, trauma, and obsessive/compulsive issues, with diverse question types, namely, factoid, diagnostic, prognostic, and preventive. We use PubMed abstracts as the primary source for QA. We develop a rigorous pipeline for LLM-based identification of information from abstracts based on various selection criteria and converting it into QA pairs. Further, valid QA pairs are extracted based on post-hoc validation criteria. Overall, our MHQA dataset consists of 2,475 expert-verified gold standard instances called MHQA-gold and ~56.1k pairs pseudo labeled using external medical references. We report F1 scores on different LLMs along with few-shot and supervised fine-tuning experiments, further discussing the insights for the scores.

心理健康问答系统大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。