构建真实医疗场景的评测基准,让大模型回答更贴近实际临床需求。
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
- 基于真实医疗场景设计多轮对话数据集,覆盖临床、健康与专业三类查询。
- 自动评分框架生成超20万条细粒度评分标准,与专家评估一致性达91.8%。
- 适合评估医疗大模型在复杂、模糊问题中的表现,尤其关注安全与准确性。
尽管大型语言模型在标准化医学考试中表现优异,但高分常无法转化为真实医疗咨询中高质量的回答。现有评估高度依赖选择题,难以捕捉真实用户提问中常见的非结构化、模糊和长尾特征。为此,我们提出 QuarkMedBench,一个面向真实医疗场景的生态有效评测基准。该数据集涵盖临床护理、健康保健与专业咨询,包含20,821个单轮查询和3,853个多轮会话。为客观评估开放式回答,我们设计了一种自动化评分框架,结合多模型共识与基于证据的检索,动态生成220,617条细粒度评分标准(平均每查询约9.8条)。评估过程中,通过层级加权与安全约束,结构化量化医疗准确性、关键点覆盖与风险拦截能力,显著降低人工评分成本与主观性。实验表明,生成的评分标准与临床专家盲审结果具有91.8%的一致性,验证了其医学可靠性。关键发现:主流模型在真实临床细节处理上表现差异显著,暴露出传统考试指标的局限性。最终,QuarkMedBench 建立了可复现、严谨的复杂健康问题评估标准,并支持动态知识更新以避免过时。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to capture the unstructured, ambiguous, and long-tail complexities inherent in genuine user inquiries. To bridge this gap, we introduce QuarkMedBench, an ecologically valid benchmark tailored for real-world medical LLM assessment. We compiled a massive dataset spanning Clinical Care, Wellness Health, and Professional Inquiry, comprising 20,821 single-turn queries and 3,853 multi-turn sessions. To objectively evaluate open-ended answers, we propose an automated scoring framework that integrates multi-model consensus with evidence-based retrieval to dynamically generate 220,617 fine-grained scoring rubrics (~9.8 per query). During evaluation, hierarchical weighting and safety constraints structurally quantify medical accuracy, key-point coverage, and risk interception, effectively mitigating the high costs and subjectivity of human grading. Experimental results demonstrate that the generated rubrics achieve a 91.8% concordance rate with clinical expert blind audits, establishing highly dependable medical reliability. Crucially, baseline evaluations on this benchmark reveal significant performance disparities among state-of-the-art models when navigating real-world clinical nuances, highlighting the limitations of conventional exam-based metrics. Ultimately, QuarkMedBench establishes a rigorous, reproducible yardstick for measuring LLM performance on complex health issues, while its framework inherently supports dynamic knowledge updates to prevent benchmark obsolescence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。