用知识图谱评估大模型在心理健康领域的认知能力
MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

- 基于PrimeKG构建九类任务,测试实体识别与关系推理
- 顶尖模型在实体识别上接近满分,但两跳推理仍表现不佳
- 输出格式可靠性影响评测结果,适合研究临床知识应用
大型语言模型(LLM)在心理健康领域应用日益广泛,但其对相关生物医学知识的掌握程度及在临床关键结构化判断中的可靠性仍不明确。本文提出一个基于知识图谱(KG)的基准测试MHGraphBench,用于评估LLMs在心理健康实体识别、关系判断和两跳推理方面的能力。该基准源自PrimeKG,包含九个任务类别,答案由知识图谱支持,并设置受控的负向选项。在15个闭源与开源模型上的实验显示,领先模型在实体分类和小规模关系分类子集上表现接近上限,但在关系预测和两跳推理任务中仍存在明显困难。此外,短知识图谱片段对部分模型有帮助,却会降低其他模型的表现。同时,在受限的多选环境下,输出格式的可靠性显著影响测评结果,凸显响应有效性在基准评估中的关键作用。因此,MHGraphBench应被理解为在受限多选界面下评估模型与经过筛选的心理健康知识片段的一致性,而非对实际临床安全性的直接衡量。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledge and how reliably they apply it to clinically salient structured judgments. Here, we present a knowledge-graph (KG)-grounded benchmark for assessing LLMs on mental-health entity recognition, relation judgment, and two-hop reasoning. The benchmark is derived from PrimeKG and comprises nine task families with KG-supported answers and controlled negative options. Experiments across 15 closed- and open-source LLMs reveal a persistent recognition-to-judgment gap: leading models achieve near-ceiling performance on entity typing and on the small relation-typing subset, yet they still struggle with relation prediction and two-hop reasoning. Additionally, short KG-derived snippets benefit some models but degrade performance for others. Moreover, output-format reliability can substantially influence measured performance under constrained multiple-choice settings, highlighting the critical role of response validity in benchmark-based evaluation. MHGraphBench should therefore be interpreted as evaluating agreement with a curated mental-health slice of PrimeKG under a constrained multiple-choice interface, rather than as a direct assessment of real-world clinical safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。