解析视觉语言模型强化微调的收敛与泛化机制,揭示奖励分解原理。
Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization
- 构建工具增强马尔可夫决策过程框架,形式化多步推理任务
- 证明复合奖励下算法以 $O(1/ ext{sqrt}{T})$ 速率收敛
- 给出奖励分解理论依据,解释小样本训练为何能跨域泛化
基于可验证奖励的强化微调(RLVR)已成为赋予大视觉语言模型(LVLM)工具使用和多步推理等代理能力的关键方法。尽管在视觉代理强化微调(Visual-ARFT)中取得显著成效,其理论基础仍不清晰。本文针对两个核心问题:(i)复合可验证奖励(格式合规、答案准确、工具可执行)如何影响组相对策略优化(GRPO)的收敛性;(ii)为何在少量工具增强任务上训练的模型能迁移到分布外领域?提出工具增强马尔可夫决策过程(TA-MDP)框架,形式化具有有限深度工具调用的多模态代理决策。主要成果包括:(1)证明在复合奖励下GRPO以 $O(1/\sqrt{T})$ 的速率收敛,明确依赖于奖励组件数和组大小;(2)提出奖励分解定理,量化分项优化与联合优化的次优差距,揭示奖励分解的适用条件;(3)建立工具增强策略的PAC-Bayes泛化界,解释Visual-ARFT中观察到的强分布外迁移能力。
原文摘要 · Abstract (English)
Reinforcement fine-tuning with verifiable rewards (RLVR) has emerged as a powerful paradigm for equipping large vision-language models (LVLMs) with agentic capabilities such as tool use and multi-step reasoning. Despite striking empirical successes, most notably Visual Agentic Reinforcement Fine-Tuning (Visual-ARFT), the theoretical underpinnings of this paradigm remain poorly understood. In particular, two critical questions lack rigorous answers: (i)~how does the composite structure of verifiable rewards (format compliance, answer accuracy, tool executability) affect the convergence of Group Relative Policy Optimization (GRPO), and (ii)~why does training on a small set of tool-augmented tasks transfer to out-of-distribution domains? We address these gaps by introducing the \emph{Tool-Augmented Markov Decision Process} (TA-MDP), a formal framework that models multimodal agentic decision-making with bounded-depth tool calls. Within this framework, we establish three main results. First, we prove that GRPO under composite verifiable rewards converges to a first-order stationary point at rate $O(1/\sqrt{T})$ with explicit dependence on the number of reward components and group size (\textbf{Theorem~1}). Second, we derive a \emph{Reward Decomposition Theorem} that bounds the sub-optimality gap between decomposed per-component optimization and joint optimization, providing a precise characterization of when reward decomposition is beneficial (\textbf{Theorem~2}). Third, we establish a PAC-Bayes generalization bound for tool-augmented policies that explains the strong out-of-distribution transfer observed in Visual-ARFT (\textbf{Theorem~3}).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。