让AI自控视觉工具调用,提速降耗还更准
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools

- 训练时让模型学会自我调节视觉工具使用时机
- 推理速度提升显著,关键任务准确率达新高
- 适合追求高效多模态推理的开发者与研究者
近期多模态大模型在细粒度视觉感知任务中取得进展,依赖频繁调用外部视觉工具和重复图像编码,导致计算开销大、推理延迟高。为此,我们提出超越视觉(BEE)——一种基于自我调节能力的隐式视觉工具范式。BEE将视觉工具调用行为纳入训练目标,使模型能自适应平衡内部知识与隐式工具,避免冗余调用,大幅降低推理延迟。具体包括两阶段训练:(1) 结构化思维链监督微调,构建带结构化工具槽和混合调用状态的思维链轨迹,激活模型的隐式工具表征与切换能力;(2) 自我调节奖励对齐,引入净工具收益(NTG)量化模糊认知边界下的冗余调用问题,设计惩罚无效依赖的奖励机制,促使模型仅在内部知识不足时才调用工具。BEE在细粒度视觉感知任务上达到顶尖性能,通用推理表现仍具竞争力,并实现显著推理效率提升。
原文摘要 · Abstract (English)
Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing various visual tool operations. However, this paradigm relies heavily on frequent external tool calls and repeated image re-encoding, which leads to substantial computational overhead and inference latency. To address these issues, we propose Beyond the Eye (BEE), a novel implicit visual tool paradigm centered on self-regulated capability. BEE directly incorporates visual tool invocation behaviors into the training objective and encourages the model to develop a self-regulated invocation mechanism. This design enables the model to adaptively balance internal knowledge and implicit tools, avoiding redundant tool usage while substantially reducing inference latency. Specifically, BEE involves a two-stage training process: (1) Formalized Chain-of-Thought (CoT) Supervised Fine-tuning (SFT). We construct CoT trajectories with structured tool slots and mixed invocation states. This stage activates the model's implicit tool representations and adaptive switching capability. (2) Self-regulated Reward-Driven Alignment. To address redundant tool usage caused by ambiguous cognitive boundaries, we first introduce the Net Tool Gain (NTG) metric to quantify this phenomenon. Based on this observation, we further propose a self-regulated reward mechanism. This mechanism penalizes ineffective tool dependency and encourages the model to perform knowledge routing, ensuring that implicit tools are invoked only when the model's internal knowledge is insufficient. BEE achieves state-of-the-art performance in fine-grained visual perception while remaining competitive in general reasoning tasks and achieving substantial gains in inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。