构建医疗AI推理评估基准,检验模型处理病历中的多步决策能力
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
- 基于真实病历数据生成包含阈值判断、时间聚合等复杂逻辑的任务
- 测试显示模型在数据聚合和阈值推理上仍有28%至64%的错误率
- 适合关注医疗AI可靠性与临床部署风险的研究者与开发者
可靠的临床决策支持需要具备在结构化电子健康记录(EHRs)上进行安全、多步骤推理的医疗AI代理。尽管大语言模型(LLMs)在医疗领域展现出潜力,但现有基准未能充分评估其在涉及阈值判断、时间聚合和条件逻辑的动作类任务上的表现。我们提出ART——一个面向医疗AI代理的动作式推理任务基准,通过挖掘真实世界EHR数据,构建针对已知推理弱点的挑战性任务。通过对现有基准的分析,我们识别出三大典型错误类型:检索失败、聚合错误和条件逻辑误判。我们的四阶段流程——场景识别、任务生成、质量审计与评估——生成了多样化且临床验证的任务。对GPT-4o-mini和Claude 3.5 Sonnet在600个任务上的评估表明,经过提示优化后检索准确率接近完美,但在聚合(28–64%)和阈值推理(32–38%)方面仍存在显著差距。ART揭示了动作导向的EHR推理中的失效模式,推动更可靠的临床智能体发展,为减轻认知负荷与行政负担、提升高需求医疗环境下的医护支持能力奠定基础。
原文摘要 · Abstract (English)
Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal aggregation, and conditional logic. We introduce ART, an Action-based Reasoning clinical Task benchmark for medical AI agents, which mines real-world EHR data to create challenging tasks targeting known reasoning weaknesses. Through analysis of existing benchmarks, we identify three dominant error categories: retrieval failures, aggregation errors, and conditional logic misjudgments. Our four-stage pipeline -- scenario identification, task generation, quality audit, and evaluation -- produces diverse, clinically validated tasks grounded in real patient data. Evaluating GPT-4o-mini and Claude 3.5 Sonnet on 600 tasks shows near-perfect retrieval after prompt refinement, but substantial gaps in aggregation (28--64%) and threshold reasoning (32--38%). By exposing failure modes in action-oriented EHR reasoning, ART advances toward more reliable clinical agents, an essential step for AI systems that reduce cognitive load and administrative burden, supporting workforce capacity in high-demand care settings
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。