首个中文医疗安全评测基准,专测大模型健康问答可靠性。
CHBench: A Chinese Dataset for Evaluating Health in Large Language Models
- 构建涵盖身心健康的6493条心理与2999条生理问题数据集
- 四款主流中文大模型在安全准确率上均存在明显短板
- 适合关注医疗AI安全、中文NLP评估的研究者使用
随着大语言模型的快速发展,评估其在健康相关问题上的表现变得愈发重要。在真实应用场景中,错误信息可能对寻求医疗建议的人造成严重后果,因此必须高度重视模型的安全性与可信度。本文提出CHBench,首个面向中文医疗场景的安全导向基准,用于评估大模型在多样化情境下理解与应对身心问题的能力。该基准包含6,493条心理健康和2,999条生理健康条目,覆盖广泛主题。对四款主流中文大模型的全面评估揭示了其在提供安全、准确健康信息方面存在显著不足,凸显该领域亟需进一步发展。代码已开源:https://github.com/TracyGuo2001/CHBench。
原文摘要 · Abstract (English)
With the rapid development of large language models (LLMs), assessing their performance on health-related inquiries has become increasingly essential. The use of these models in real-world contexts-where misinformation can lead to serious consequences for individuals seeking medical advice and support-necessitates a rigorous focus on safety and trustworthiness. In this work, we introduce CHBench, the first comprehensive safety-oriented Chinese health-related benchmark designed to evaluate LLMs' capabilities in understanding and addressing physical and mental health issues with a safety perspective across diverse scenarios. CHBench comprises 6,493 entries on mental health and 2,999 entries on physical health, spanning a wide range of topics. Our extensive evaluations of four popular Chinese LLMs highlight significant gaps in their capacity to deliver safe and accurate health information, underscoring the urgent need for further advancements in this critical domain. The code is available at https://github.com/TracyGuo2001/CHBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。