评测开源医学大模型的可信度,发现性能与安全存权衡
Truth, Trust, and Trouble: Medical AI on the Edge
- 构建千级医疗问答数据集,评估模型诚实、有用、无害性
- AlpaCare-13B准确率最高达91.7%,但复杂问题帮助性下降
- 领域微调提升安全性,少样本提示可增准率至85%
大型语言模型在数字健康领域具有变革潜力,可实现自动化医疗问答。然而,确保其在事实准确性、实用性与安全性方面符合行业标准仍具挑战,尤其对开源模型而言。本文提出一个严格的基准测试框架,使用超过1,000个健康相关问题的数据集,评估Mistral-7B、BioMistral-7B-DARE和AlpaCare-13B等模型在诚实性、有用性和无害性方面的表现。结果显示,各模型在事实可靠性与安全性之间存在权衡:AlpaCare-13B达到最高准确率(91.7%)和无害性评分(0.92),而领域特定微调的BioMistral-7B-DARE虽规模较小,但安全性达0.90;少样本提示使准确率从78%提升至85%,所有模型在复杂查询上的有用性均下降,揭示临床问答任务仍面临挑战。
原文摘要 · Abstract (English)
Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. However, ensuring these models meet critical industry standards for factual accuracy, usefulness, and safety remains a challenge, especially for open-source solutions. We present a rigorous benchmarking framework using a dataset of over 1,000 health questions. We assess model performance across honesty, helpfulness, and harmlessness. Our results highlight trade-offs between factual reliability and safety among evaluated models -- Mistral-7B, BioMistral-7B-DARE, and AlpaCare-13B. AlpaCare-13B achieves the highest accuracy (91.7%) and harmlessness (0.92), while domain-specific tuning in BioMistral-7B-DARE boosts safety (0.90) despite its smaller scale. Few-shot prompting improves accuracy from 78% to 85%, and all models show reduced helpfulness on complex queries, highlighting ongoing challenges in clinical QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。