用权威临床指南重构医疗AI评估标准,提升模型可靠性与全球适用性。
Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench
- 以系统评价和GRADE评级的临床指南替代专家意见作为评分依据
- 提出证据加权评分与上下文优先逻辑,减少偏见并增强临床可信度
- 适合关注医疗AI公平性、临床落地与全球适配的研究者与开发者
HealthBench 是一项旨在评估 AI 在医疗语言任务中能力的基准(Arora 等,2025),通过医生设计的对话和透明评分规则推动了医学语言模型的评测进展。然而,其依赖专家意见而非高等级临床证据,可能固化地区偏见和个体医生习惯,并因自动化评分系统引入额外偏差。这些缺陷在低收入与中等收入国家尤为突出,如热带病覆盖不足、指南区域性差异等问题普遍存在。非洲等地区面临数据稀缺、基础设施薄弱及监管体系不健全等挑战,亟需更具全球普适性和公平性的评测基准。为此,我们建议将奖励函数锚定于版本可控的临床实践指南(CPGs),整合系统评价与GRADE证据分级。本路线图提出基于指南-评分规则映射的‘证据稳健’强化学习,包括证据加权评分、上下文优先逻辑,并结合伦理考量与延迟结果反馈。通过将奖励机制扎根于严格审核的临床指南,同时保留 HealthBench 的透明性与医生参与,旨在培育语言流畅且临床可信、伦理合规、全球适用的医学语言模型。
原文摘要 · Abstract (English)
HealthBench, a benchmark designed to measure the capabilities of AI systems for health better (Arora et al., 2025), has advanced medical language model evaluation through physician-crafted dialogues and transparent rubrics. However, its reliance on expert opinion, rather than high-tier clinical evidence, risks codifying regional biases and individual clinician idiosyncrasies, further compounded by potential biases in automated grading systems. These limitations are particularly magnified in low- and middle-income settings, where issues like sparse neglected tropical disease coverage and region-specific guideline mismatches are prevalent. The unique challenges of the African context, including data scarcity, inadequate infrastructure, and nascent regulatory frameworks, underscore the urgent need for more globally relevant and equitable benchmarks. To address these shortcomings, we propose anchoring reward functions in version-controlled Clinical Practice Guidelines (CPGs) that incorporate systematic reviews and GRADE evidence ratings. Our roadmap outlines "evidence-robust" reinforcement learning via rubric-to-guideline linkage, evidence-weighted scoring, and contextual override logic, complemented by a focus on ethical considerations and the integration of delayed outcome feedback. By re-grounding rewards in rigorously vetted CPGs, while preserving HealthBench's transparency and physician engagement, we aim to foster medical language models that are not only linguistically polished but also clinically trustworthy, ethically sound, and globally relevant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。