arXiv:2604.19281cs.HCcs.AI2026-04中稿 · the Ninth Annual A…

提出新评估框架VB-Score,发现医疗大模型在慢性病领域存在严重性能偏差。

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

论文配图:Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
图 1 · 摘自论文原文
  • 按实体识别、语义相似度等四维度分别评估医疗问答模型
  • 三款主流模型在老年及少数族裔相关慢性病上性能低13.8%以上
  • 提示工程无法弥补模型架构缺陷,提醒仅靠语义匹配不安全

大型语言模型(LLMs)在辅助患者解答医学问题方面日益普及。然而,当前多数评估方法仅衡量答案与语义的匹配程度,无法真实反映模型的医学准确性或健康公平性风险。为此,我们提出一种名为VB-Score(基于验证的评分)的新评估框架,对医疗问答模型的实体识别、语义相似度、事实一致性和结构化信息完整性四个组件进行独立评估。我们在48个来自高质量权威来源的公共卫生主题上,对三款知名且广泛使用的LLMs进行了严格评测。分析发现,模型在语义与实体准确性之间存在显著差异。所有模型在各项标准下均表现严重不足。尤其在与老年人群和少数族裔相关的慢性病议题上,模型平均性能比整体水平低13.8%,表明存在所谓的基于病症的算法歧视。研究还显示,仅靠提示工程无法弥补模型在医学实体提取上的根本性局限,质疑了仅以语义匹配作为医疗AI安全评价标准的充分性。

原文摘要 · Abstract (English)

The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the measures currently used to evaluate the performance of these models in this context only measure how closely a model's answers match semantically, and therefore do not provide a true indication of the model's medical accuracy or of the health equity risks associated with it. To address these shortcomings, we present a new evaluation framework for medical question answering called VB-Score (Verification-Based Score) that provides a separate evaluation of the four components of entity recognition, semantic similarity, factual consistency, and structured information completeness for medical question-answering models. We perform rigorous reviews of the performance of three well-known and widely used LLMs on 48 public health-related topics taken from high-quality, authoritative information sources. Based on our analyses, we discover a major discrepancy between the models' semantic and entity accuracy. Our assessments of the performance of all three models show that each of them has almost uniformly severe performance failures when evaluated against our criteria. Our findings indicate alarming performance disparities across various public health topics, with most of the models exhibiting 13.8% lower performance (compared to an overall average) for all the public health topics that relate to chronic conditions that occur in older and minority populations, which indicates the existence of what's known as condition-based algorithmic discrimination. Our findings also demonstrate that prompt engineering alone does not compensate for basic architectural limitations on how these models perform in extracting medical entities and raise the question of whether semantic evaluation alone is a sufficient measure of medical AI safety.

医疗AI评估框架算法偏见大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。