arXiv:2607.25186cs.CL2026-07

首个覆盖心血管全周期的真实世界多任务评测基准,评估大模型临床实用性。

CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

论文配图:CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
图 1 · 摘自论文原文
  • 构建涵盖心血管全周期的2263个真实病例任务,由16名心内科医生标注。
  • GPT-5.4综合表现最优,但伦理与心电图解读得分最低仅17.34。
  • 发现模型在沟通、急救等场景中关键信息遗漏严重,适合临床研究者参考。

背景:当前医疗大模型评测多聚焦于考试知识或孤立任务,难以反映心血管诊疗中纵向、多模态和高安全性的实际流程。目的:开发CardioBench——一个覆盖心血管照护全流程的真实世界基准,并评估大模型在临床维度与专科任务中的表现。方法:CardioBench包含来自13个任务数据集的2,263个条目,源自去标识化的心血管病历与检查数据。16名心脏病学专家完成标注与参考构建,经两名资深心脏病专家交叉评审。七款大模型在标准化零样本设置下生成15,841条输出。开放性任务采用关键点覆盖率与整体临床质量评估,CardioEthics则以准确性评分。结果:GPT-5.4在宏平均(62.55)和项加权均值(62.19)上最高,其次为Gemini 3.1 Pro(59.95)和Qwen 3.6 27B(59.72),并在所有三个维度排名第一。CardioAuxReport表现最佳(86.38),而CardioECGRead(17.25)与CardioEthics(17.34)最低。整体临床质量与关键点覆盖率差距最大的是CardioComm(52.71)、CardioEmergRescue(52.05)和CardioTreatPlan(48.80)。结论:据我们所知,CardioBench是目前最大、最全面的真实世界多任务基准,覆盖了迄今为止报道最广泛的临床真实心血管场景,为识别模型优势、关键缺失及未来发展方向提供了严谨框架。

原文摘要 · Abstract (English)

Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop CardioBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: CardioBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, CardioBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.

心内科大模型评测临床验证真实世界数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。