arXiv:2510.18424cs.AI2025-10EMNLP被引 5

用搜索与反馈提升医学视觉推理模型的准确性与一致性

Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents

  • 结合视觉引导与蒙特卡洛树搜索,增强模型推理过程
  • 在多个医学VQA数据集上超越现有方法,减少幻觉与逻辑错误
  • 适合医疗AI研发者及需要高可靠性视觉推理的应用场景

视觉语言模型(VLMs)在医学推理任务中表现优异,但存在幻觉、描述模糊、逻辑不一致和定位能力差等问题。为此,我们提出一种名为医学视觉推理代理(Med-VRAgent)的代理框架,基于视觉引导与自奖励机制,并结合蒙特卡洛树搜索(MCTS)。通过将视觉引导与树搜索相结合,显著提升了VLMs的医学视觉推理能力。我们利用Med-VRAgent生成的推理轨迹作为反馈,采用近端策略优化(PPO)目标对VLMs进行微调,进一步提升性能。在多个医学视觉问答(VQA)基准测试中,该方法均优于现有方法。

原文摘要 · Abstract (English)

Visual Language Models (VLMs) achieve promising results in medical reasoning but struggle with hallucinations, vague descriptions, inconsistent logic and poor localization. To address this, we propose a agent framework named Medical Visual Reasoning Agent (\textbf{Med-VRAgent}). The approach is based on Visual Guidance and Self-Reward paradigms and Monte Carlo Tree Search (MCTS). By combining the Visual Guidance with tree search, Med-VRAgent improves the medical visual reasoning capabilities of VLMs. We use the trajectories collected by Med-VRAgent as feedback to further improve the performance by fine-tuning the VLMs with the proximal policy optimization (PPO) objective. Experiments on multiple medical VQA benchmarks demonstrate that our method outperforms existing approaches.

医学视觉推理视觉语言模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。