让医疗研究智能体自我进化,提升医学影像分析能力
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

- 用阶段评分体系指导智能体持续优化
- 90亿参数模型在基准测试中得分超79.6,超越Claude Opus
- 适合需要长期迭代的医疗AI研发团队使用
长周期智能体正开始自动化生成代码、报告和研究成果的完整流程。医学影像分析流程多阶段且依赖数据,但专家行为轨迹稀缺且难以共享。结构化基准可通过阶段级评分定位失败,但标准后训练会丢弃这些诊断信息。本文提出基准即教师(BaT)系统,一种递归自改进的后训练框架。其包含两个联动组件:异步阶段银行数据流水线与双层课程强化学习(BiCuRL)。阶段银行在策略更新循环外合成内容隔离的训练状态;BiCuRL利用固定保留评估集选择下一阶段课程,以任务评分验证生成结果,通过GRPO更新策略,并将候选检查点返回评估。在AutoMedBench-Lite上,BaT-4B和BaT-9B的综合得分超过其Qwen Instruct基线两倍以上。其中,BaT-9B智能体达到79.6的综合得分,优于使用Claude Code的Claude Opus 4.6(77.5)。
原文摘要 · Abstract (English)
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。