M3Builder用四个智能体自动完成医学影像机器学习全流程。
M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging
- 四类专用智能体协作,从数据处理到模型训练全程自动化。
- 在14个数据集上实现94.29%任务成功率,使用Claude-3.7-Sonnet作为核心。
- 专为医学影像设计,适合医疗AI研发人员和自动化系统研究者。
智能体式AI系统因能自主执行复杂任务而受到广泛关注,但其对预设工具的依赖限制了在医学领域的应用,而该领域需定制化模型训练。本文提出三项贡献:(i) 提出M3Builder,一种新型多智能体系统,用于自动化医学影像机器学习。系统包含四个专业智能体,协同完成从数据处理、环境配置到自包含自动调试与模型训练的多步骤流程。它们运行于一个结构化的医学影像机器学习工作区,通过自然语言描述数据集、训练代码和交互工具实现无缝沟通与任务执行。(ii) 构建M3Bench基准,涵盖14个训练数据集、五种解剖部位、三种成像模态,覆盖2D与3D数据,评估自动化医学影像机器学习进展。(iii) 在七种主流大语言模型(如Claude系列、GPT-4o、DeepSeek-V3)作为智能体核心的实验中,相较于现有设计,M3Builder在医学影像任务上表现更优,采用Claude-3.7-Sonnet时达成94.29%的成功率,展现出迈向完全自动化机器学习的巨大潜力。
原文摘要 · Abstract (English)
Agentic AI systems have gained significant attention for their ability to autonomously perform complex tasks. However, their reliance on well-prepared tools limits their applicability in the medical domain, which requires to train specialized models. In this paper, we make three contributions: (i) We present M3Builder, a novel multi-agent system designed to automate machine learning (ML) in medical imaging. At its core, M3Builder employs four specialized agents that collaborate to tackle complex, multi-step medical ML workflows, from automated data processing and environment configuration to self-contained auto debugging and model training. These agents operate within a medical imaging ML workspace, a structured environment designed to provide agents with free-text descriptions of datasets, training codes, and interaction tools, enabling seamless communication and task execution. (ii) To evaluate progress in automated medical imaging ML, we propose M3Bench, a benchmark comprising four general tasks on 14 training datasets, across five anatomies and three imaging modalities, covering both 2D and 3D data. (iii) We experiment with seven state-of-the-art large language models serving as agent cores for our system, such as Claude series, GPT-4o, and DeepSeek-V3. Compared to existing ML agentic designs, M3Builder shows superior performance on completing ML tasks in medical imaging, achieving a 94.29% success rate using Claude-3.7-Sonnet as the agent core, showing huge potential towards fully automated machine learning in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。