arXiv:2602.10619cs.CVcs.AI2026-02被引 1

医学影像模型通过增强视觉感知与推理能力,显著提升强化微调效果。

Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation

  • 引入先验知识注入与感知驱动策略,强化模型对医学图像的理解
  • 设计医疗导向的奖励机制与行为模仿,使模型决策更符合临床逻辑
  • 在多个数据集上超越传统微调和强化微调基线,适合高风险医疗场景

尽管近期强化微调(RFT)在大语言模型后训练中取得进展,其在跨模态、以视觉为中心的领域仍研究不足,尤其在医学影像领域,高效表现需兼具强视觉感知与结构化推理。本文提出专为医学领域设计的视觉强化微调框架VRFT-Aug,通过引入先验知识注入、感知驱动的策略优化、医疗导向的奖励塑造及行为模仿等训练策略,共同增强感知与推理能力,稳定并提升RFT过程。在多个医学数据集上的大量实验表明,该方法持续优于标准监督微调与RFT基线。此外,我们提供了可推广至其他医学图像任务的实证洞察与实用训练启发。本工作旨在为高风险医疗应用中构建可靠、具备推理能力的模型提供可操作指导与新思路。

原文摘要 · Abstract (English)

While recent advances in Reinforcement Fine-Tuning (RFT) have shown that rule-based reward schemes can enable effective post-training for large language models, their extension to cross-modal, vision-centric domains remains largely underexplored. This limitation is especially pronounced in the medical imaging domain, where effective performance requires both robust visual perception and structured reasoning. In this work, we address this gap by proposing VRFT-Aug, a visual reinforcement fine-tuning framework tailored for the medical domain. VRFT-Aug introduces a series of training strategies designed to augment both perception and reasoning, including prior knowledge injection, perception-driven policy refinement, medically informed reward shaping, and behavioral imitation. Together, these methods aim to stabilize and improve the RFT process. Through extensive experiments across multiple medical datasets, we show that our approaches consistently outperform both standard supervised fine-tuning and RFT baselines. Moreover, we provide empirically grounded insights and practical training heuristics that can be generalized to other medical image tasks. We hope this work contributes actionable guidance and fresh inspiration for the ongoing effort to develop reliable, reasoning-capable models for high-stakes medical applications.

医学影像强化微调视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。