arXiv:2604.10696cs.AIcs.CV2026-04被引 4

Camyla实现医学影像分割的全自动科研,自动生成论文与模型。

Camyla: Scaling Autonomous Research in Medical Image Segmentation

论文配图:Camyla: Scaling Autonomous Research in Medical Image Segmentation
图 1 · 摘自论文原文
  • 用质量加权分支探索、分层反思记忆和发散诊断反馈三机制应对自动实验挑战。
  • 28天内生成超2700个新模型和40篇完整论文,24个数据集超越最强基线。
  • 适合关注自动化科研、医学影像分析的研究者与开发者。

我们提出Camyla,一个在医学图像分割领域实现完全自主研究的系统。Camyla将原始数据集转化为基于文献的研究方案、可执行实验及完整论文,全程无需人工干预。长期自主实验面临三大挑战:搜索精力偏离有效方向、早期试验知识随上下文累积而退化、失败后恢复陷入重复微调。为此,系统融合三项耦合机制:质量加权分支探索用于分配多方案实验资源,分层反思记忆用于在多粒度上保留并压缩跨试验知识,发散诊断反馈用于在性能不佳后多样化修复策略。系统在CamylaBench上评估,该基准由31个2025年发表论文构建,无污染数据,采用严格零干预协议,在8张GPU集群上完成两次独立运行,总计28天。两次运行中,Camyla生成超过2700个新型模型实现与40篇完整论文,分别在22和18个数据集上超越14种成熟架构(包括nnU-Net)的单数据集最强基线(并集:24/31)。资深人类评审认为生成论文达到当代医学影像期刊T1/T2水平。相比自动化基线,Camyla在整体分割性能上优于AutoML与NAS系统,并在任务完成率与超越基线频率上超过六个开放研究代理。结果表明,医学图像分割领域的规模化自主研究已可实现。

原文摘要 · Abstract (English)

We present Camyla, a system for fully autonomous research within the scientific domain of medical image segmentation. Camyla transforms raw datasets into literature-grounded research proposals, executable experiments, and complete manuscripts without human intervention. Autonomous experimentation over long horizons poses three interrelated challenges: search effort drifts toward unpromising directions, knowledge from earlier trials degrades as context accumulates, and recovery from failures collapses into repetitive incremental fixes. To address these challenges, the system combines three coupled mechanisms: Quality-Weighted Branch Exploration for allocating effort across competing proposals, Layered Reflective Memory for retaining and compressing cross-trial knowledge at multiple granularities, and Divergent Diagnostic Feedback for diversifying recovery after underperforming trials. The system is evaluated on CamylaBench, a contamination-free benchmark of 31 datasets constructed exclusively from 2025 publications, under a strict zero-intervention protocol across two independent runs within a total of 28 days on an 8-GPU cluster. Across the two runs, Camyla generates more than 2,700 novel model implementations and 40 complete manuscripts, and surpasses the strongest per-dataset baseline selected from 14 established architectures, including nnU-Net, on 22 and 18 of 31 datasets under identical training budgets, respectively (union: 24/31). Senior human reviewers score the generated manuscripts at the T1/T2 boundary of contemporary medical imaging journals. Relative to automated baselines, Camyla outperforms AutoML and NAS systems on aggregate segmentation performance and exceeds six open-ended research agents on both task completion and baseline-surpassing frequency. These results suggest that domain-scale autonomous research is achievable in medical image segmentation.

医学影像自动化科研自生成论文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。