arXiv:2607.10522cs.CVcs.AI2026-07

提出可自主开发医疗影像模型的多智能体系统,提升效率与可审计性。

Towards Autonomous and Auditable Medical Imaging Model Development

论文配图:Towards Autonomous and Auditable Medical Imaging Model Development
图 1 · 摘自论文原文
  • 基于数据条件的方法规划,生成可并行执行的模型路径
  • 两阶段优化确保验证流程、指标计算与预测结果严格合规
  • 在20项任务中超越通用ML系统,接近人类专家水平

大型语言模型代理正开始通过规划、代码执行、调试和经验反馈自动化机器学习工程。将这一能力应用于医疗影像仍具挑战,因每项任务需特定模态实验,并对验证协议与预测结果有严格要求。本文提出AMID,一种面向医疗影像模型开发的自主多智能体框架。AMID首先提出数据条件方法规划,将粗略的任务搜索空间细化为基于任务特定数据分析和可运行资源的可执行、可并行化方法路径。随后构建验证引导的两阶段优化,从广泛探索多样方法路径转向对有潜力候选者的定向利用,同时在整个优化过程中严格验证验证协议、指标计算和预测结果。在涵盖多种模态和预测类型的20个医疗影像挑战任务中,AMID优于评估过的通用机器学习工程系统,且在若干任务上达到或接近人类设计的优秀解决方案水平。结果表明,AMID可将任务特定的医疗影像模型开发从定制化手动工程转变为可代理化的高绩效、可审计工作流,适用于异构任务。

原文摘要 · Abstract (English)

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.

医疗影像多智能体自动化可审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。