arXiv:2409.07314cs.CLcs.AI2024-09被引 29

构建临床大模型评估框架,揭示知识与实操能力的显著差距。

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

  • 设计多维度评估框架,量化模型在临床任务中的真实能力
  • 发现知识检索强不等于实操能力强,存在明显执行鸿沟
  • 适合医疗AI研发者和临床部署团队参考决策

尽管大语言模型(LLMs)在标准化医学执照考试中表现超人,但这些静态基准已趋于饱和,与临床工作流的实际需求日益脱节。为弥合理论能力与实际效用之间的差距,我们提出MEDIC——一个全面的评估框架,建立临床领域大模型能力的五大领先指标。这些前置指标揭示了跨基准的能力差异,例如静态知识检索与功能执行之间的分歧,可在高成本部署前指导模型选择。除常规问答外,我们采用确定性执行协议和新型交叉审查框架(CEF),在无需参考文本的情况下量化信息保真度与幻觉率。在异构任务集上的评估暴露了关键性能权衡:我们发现显著的知识-执行差距,即静态检索能力强并不预示临床计算或SQL生成等操作任务的成功。此外,被动安全(拒绝回答)与主动安全(错误检测)之间存在分化,经过高拒绝率微调的模型往往无法可靠审核临床文档的真实性。结果表明,无单一架构在所有维度占优,强调临床模型部署需采用组合策略。本研究附带公开可访问的MEDIC排行榜(https://hf.co/spaces/m42-health/MEDIC-Benchmark)。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To bridge the gap between theoretical capability and verified utility, we introduce MEDIC, a comprehensive evaluation framework establishing leading indicators of clinical LLM competence across five dimensions. These upfront indicators reveal cross-benchmark capability gaps, such as the divergence between static knowledge retrieval and functional execution, that inform model selection before costly deployment-based evaluation. Beyond standard question-answering, we assess operational capabilities using deterministic execution protocols and a novel Cross-Examination Framework (CEF), which quantifies information fidelity and hallucination rates without reliance on reference texts. Our evaluation across a heterogeneous task suite exposes critical performance trade-offs: we identify a significant knowledge-execution gap, where proficiency in static retrieval does not predict success in operational tasks such as clinical calculation or SQL generation. Furthermore, we observe a divergence between passive safety (refusal) and active safety (error detection), revealing that models fine-tuned for high refusal rates often fail to reliably audit clinical documentation for factual accuracy. These findings demonstrate that no single architecture dominates across all dimensions, highlighting the necessity of a portfolio approach to clinical model deployment. We accompany this work with a publicly available MEDIC leaderboard at https://hf.co/spaces/m42-health/MEDIC-Benchmark.

大模型评估临床AI安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。