用工具监督强化学习,让大模型更会用视觉工具解题。
Visual Reasoning through Tool-supervised Reinforcement Learning

- 通过直接工具监督,分阶段训练模型掌握绘图、旋转等操作。
- 在复杂视觉任务上,工具调用能力显著提升,准确率更高。
- 适合需要精准视觉推理的AI研究者与开发者参考。
本文研究如何让多模态大模型有效掌握工具使用以解决复杂视觉推理任务。为此,我们提出一种新型工具监督强化学习(ToolsRL)框架,采用易于收集的直接工具监督机制。聚焦一系列简单、原生且可解释的视觉工具,包括缩放、旋转、翻转及绘制点/线。设计分阶段强化学习课程:第一阶段仅通过特定工具奖励优化,第二阶段在追求准确率的同时允许调用工具。该策略先掌握工具使用能力,再用于完成任务,避免不同任务间的优化冲突。实验表明,工具监督课程训练高效,ToolsRL在复杂视觉推理任务中展现出强大的工具使用能力。
原文摘要 · Abstract (English)
In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Models. To achieve that, we propose a novel Tool-supervised Reinforcement Learning (ToolsRL) framework, with direct tool supervision for more effective tool-use learning. We focus on a series of simple, native, and interpretable visual tools, including zoom-in, rotate, flip, and draw point/line, whose tool supervision is easy to collect. A reinforcement learning curriculum is developed, where the first stage is solely optimized by a set of well motivated tool-specific rewards, and the second stage is trained with the accuracy targeted rewards while allowing calling tools. In this way, tool calling capability is mastered before using tools to complete visual reasoning tasks, avoiding the potential optimization conflict among those heterogeneous tasks. Our experiments have shown that the tool-supervised curriculum training is efficient and ToolsRL can achieve strong tool-use capabilities for complex visual reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。