医学影像领域首个端到端自主智能体评测基准,揭示当前AI在专业场景下的严重不足。
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
- 构建20个真实医学影像竞赛挑战,覆盖多模态与任务类型
- 要求智能体独立完成数据预处理、训练与提交,限时限算力
- 现有顶尖代理多数表现不及人类专家0百分位,暴露知识与工程短板
基于大语言模型的自主编码智能体虽能处理通用软件与机器学习任务,但在复杂、特定领域的科学问题上仍显乏力。医学影像尤其具挑战性,需长期训练周期、高维数据处理及专用预处理与验证流程,而现有智能体评测体系未能充分衡量这些能力。为此,我们提出ReX-MLE,一个由20个来自高影响力医学影像竞赛的挑战构成的基准,涵盖多样模态与任务类型。不同于以往的机器学习智能体评测,ReX-MLE评估完整的端到端工作流,要求智能体在真实计算与时间约束下自主完成数据预处理、模型训练与结果提交。对主流智能体(AIDE、ML-Master、R&D-Agent)搭配不同大模型后端(GPT-5、Gemini、Claude)进行评估,发现其表现存在严重差距:多数提交结果排名位于人类专家的0百分位。失败原因主要源于领域知识缺失与工程实现限制。ReX-MLE揭示了这些瓶颈,并为开发具备领域感知能力的自主AI系统提供了基础。
原文摘要 · Abstract (English)
Autonomous coding agents built on large language models (LLMs) can now solve many general software and machine learning tasks, but they remain ineffective on complex, domain-specific scientific problems. Medical imaging is a particularly demanding domain, requiring long training cycles, high-dimensional data handling, and specialized preprocessing and validation pipelines, capabilities not fully measured in existing agent benchmarks. To address this gap, we introduce ReX-MLE, a benchmark of 20 challenges derived from high-impact medical imaging competitions spanning diverse modalities and task types. Unlike prior ML-agent benchmarks, ReX-MLE evaluates full end-to-end workflows, requiring agents to independently manage data preprocessing, model training, and submission under realistic compute and time constraints. Evaluating state-of-the-art agents (AIDE, ML-Master, R&D-Agent) with different LLM backends (GPT-5, Gemini, Claude), we observe a severe performance gap: most submissions rank in the 0th percentile compared to human experts. Failures stem from domain-knowledge and engineering limitations. ReX-MLE exposes these bottlenecks and provides a foundation for developing domain-aware autonomous AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。