arXiv:2607.26155cs.AI2026-07

构建首个面向长期临床数据的可执行分析基准,推动医疗智能体从能运行到真正确。

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

  • 以程序优先反向合成法构建可验证的临床分析任务
  • 最强模型仅56.3%任务达成严格通过,远低于执行成功率
  • 适合医疗AI、临床推理系统研究者参考

临床数据科学智能体需将异构的长期记录转化为可审计的分析结果,但现有基准多孤立地处理医学问答、表格推理或通用科研资源。我们提出CLINLENS,一个涵盖五类关联MIMIC数据(结构化电子病历、病程记录、心电图、胸片、超声心动图)的200个可执行任务基准。采用4×5的患者-时间范围与分析能力交叉分类法。程序优先的反向合成方法为每个半原始数据包配对评估私有的参考工作流,并验证所需实体、队列与时间语义及最终答案。在固定126个任务集上,24种标准化模型配置中最强者达56.3%的范围宏观严格通过率,尽管执行成功率达100%。作为参照,独立配置的编码智能体解决83/126任务,而五种适配GPT-4o-mini的生物医学系统最高仅2.9%达到范围宏观严格通过。这些结果揭示了可运行提交与正确临床分析之间的巨大差距。

原文摘要 · Abstract (English)

Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms. A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities. Program-first reverse synthesis pairs each bounded semi-raw package with an evaluator-private reference workflow and checks required artifacts, cohort and temporal semantics, and the final answer. On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS despite 100% EXECSUCCESS. For reference, a separately configured coding agent solves 83 of 126 tasks, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS. These results expose a substantial gap between runnable submissions and correct clinical analyses.

临床智能体多模态长时序基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。