arXiv:2608.21864cs.LGcs.AI2026-08

用强化学习打造可自适应的医学智能代理,提升诊断可靠性

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

  • 通过元学习与强化学习动态整合多模型专家能力
  • 在多个基准上达73%准确率,较现有模型提升约5%
  • 适合需要高可靠医疗推理的临床研究与系统开发

当前临床视觉大语言模型虽推动了数字诊断进步,但仍面临病灶噪声、模态错配、幻觉及上下文缺失等问题。现有智能体系统多依赖静态非自适应流程,缺乏复杂医学推理所需的灵活性。为此,我们提出BioMed-Agent-RL,一种融合自适应编排、策略与基于奖励的强化学习的统一医学代理。为确保可靠性,引入临床上下文感知偏好优化(CPO)、直接偏好优化(DPO)与组相对策略优化(GRPO),并结合动态熵调节。该框架采用多模态元学习,作为领域专家与人类判断的合成器,能根据临床模态(如X光)自适应调用临床定位、推理、病灶分割与领域特定合成等模型级专长,通过迭代式自适应强化学习,学会在存在误导性或矛盾视觉线索时仍坚持内在推理,即使专家建议有误亦然。在多个基准上的消融实验表明,该代理显著优于现有最先进模型,如GPT-5,准确率最高达~73%,较当代基线提升~5%。该框架为构建事实性、可靠、鲁棒且类专家的独立临床推理智能体提供了新标准。

原文摘要 · Abstract (English)

The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often endure lesion noises, modality misalignment, hallucination, and missed contextual grounding in complex clinical cases. Moreover, prevailing agent systems usually depend on static and non-adaptable pipelines and lack the versatility necessary for complex medical reasoning. To resolve these difficulties, we present BioMed-Agent-RL, a unified medical agent that incorporates adaptive orchestration, policy, and reward-based reinforcement learning (RL) models for biomedical applications. To ensure reliability, it invokes clinical context-aware preference optimization (CPO), direct preference optimization (DPO), and group relative policy optimization (GRPO) with dynamic entropy regulation. This pipeline utilizes a multimodal meta-learning approach that operates as a field-specific expert and human judgment synthesizer. The agent adaptively utilizes a set of model-level expertise, such as clinical grounding and reasoner, lesion segmenter, and field-specific synthesizer, across various clinical modalities (e.g., X-ray) by utilizing an iterative and adaptive RL approach. The agent learns to seriously synthesize misleading, conflicting vision cues and trust in inherent reasoning, while specialist advice is faulty. An intensive ablation study is conducted across multiple benchmarks, and the agent significantly outperforms existing state of the art models, such as GPT-5, attaining up to ~73% accuracy (gain of ~5%) over contemporary baselines. As a result, the framework suggests a new standard for building factual, reliable, robust, and expert-like intelligent agent systems for independent clinical reasoning.

医学智能体强化学习多模态临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。