用工具增强大模型,让细粒度图像分类更可靠。
ToolFG: Towards Well-Grounded Fine-Grained Image Classification

- 让大模型自主调用外部工具,主动获取视觉线索
- 通过强化学习优化工具与模型协同,提升分类准确率
- 适合需要高精度图像识别的科研与工业场景
细粒度图像分类(FGIC)应用广泛,备受研究关注。本文提出首个面向FGIC的工具集成多模态大模型框架ToolFG。ToolFG使多模态大模型在推理过程中能自主、灵活地调用外部工具,主动与图像交互,收集可验证的视觉线索,以更可靠、更坚实的方式区分高度相似的类别。为赋予模型这一工具使用能力,我们设计了一种基于蒙特卡洛树搜索(MCTS)引导的工具使用知识蒸馏机制,从先进的专有多模态大模型中有效挖掘与工具使用及FGIC相关的知识用于训练。此外,提出一种模型-工具协同进化机制,联合优化工具集与模型的工具使用策略,推动二者向相互适应且专精于FGIC的状态演进。大量实验验证了该框架的有效性。
原文摘要 · Abstract (English)
Fine-grained image classification (FGIC) has broad applications and has attracted significant research attention. In this paper, we explore a novel paradigm for solving FGIC by proposing \textbf{ToolFG}, the first tool-integrated MLLM-based framework tailored to FGIC. ToolFG enables MLLMs to autonomously and flexibly use external tools during the reasoning process, actively interact with images, and collect verifiable visual cues for distinguishing highly similar categories in a more \textit{reliable} and \textit{well-grounded} manner. To equip the model with such tool-use ability, we design a novel \textbf{MCTS-guided tool-use knowledge distillation mechanism}, which effectively mines tool-use- and FGIC-relevant knowledge from advanced proprietary MLLMs for model training. Furthermore, we propose a \textbf{model-tool co-evolution mechanism} that jointly refines the toolset and the model's tool-use policy, driving them toward a mutually adapted and FGIC-specialized state. Extensive experiments demonstrate the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。