强化微调能提升视觉任务表现,但复杂任务才受益
On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
- 用强化微调优化视觉模型,比传统微调更有效
- 样本少时强化微调优势明显,复杂任务表现更好
- 过度思考会拖累简单任务,需合理控制推理深度
强化微调(RFT)在提升大语言模型推理能力方面已被证明极具价值。研究人员正尝试将其应用于多模态大模型(MLLMs),以增强视觉理解能力。然而,相关研究仍处于早期阶段,尚未系统评估RFT在视觉任务中的适用性。本文通过实验分析和观察,探究RFT在视觉任务中的潜力与局限。定量比较显示,RFT在多数视觉任务中优于监督微调(SFT),尤其在训练样本有限时。为进一步检验优势是否源于推理过程,我们设计了一种鼓励模型“更多思考”的新奖励机制。结果表明,增加推理深度对复杂任务有益,但对简单任务反而有害。本研究为该方向的快速发展提供了关键洞察。
原文摘要 · Abstract (English)
Reinforcement Fine-Tuning (RFT) is proved to be greatly valuable for enhancing the reasoning ability of LLMs. Researchers have been starting to apply RFT to MLLMs, hoping it will also enhance the capabilities of visual understanding. However, these works are at a very early stage and have not examined how suitable RFT actually is for visual tasks. In this work, we endeavor to understand the suitabilities and limitations of RFT for visual tasks, through experimental analysis and observations. We start by quantitative comparisons on various tasks, which shows RFT is generally better than SFT on visual tasks. %especially when the number of training samples are limited. To check whether such advantages are brought up by the reasoning process, we design a new reward that encourages the model to ``think'' more, whose results show more thinking can be beneficial for complicated tasks but harmful for simple tasks. We hope this study can provide more insight for the rapid advancements on this topic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。