arXiv:2601.02825cs.CV2026-01

让大模型像人一样用简笔画式推理,大幅降低计算开销。

SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models

  • 将长推理转为简洁目标导向的简笔画式推理
  • 推理令牌成本降低64%以上,答案准确率不变
  • 适合追求高效推理的多模态应用开发者

尽管大规模多模态模型的逐步推理在实践中表现良好,但漫长的推理过程不可避免地带来显著计算开销,包括更高的令牌消耗和更长响应时间,损害了推理效率。相比之下,人类常采用简笔画式推理:一种简洁、目标导向的认知过程,优先关注关键信息,实现高效问题求解。受此认知效率启发,我们提出SketchThinker-R1,激励大规模多模态模型具备简笔画式推理能力。方法包含三个阶段:首先在草图模式冷启动阶段,将标准长推理转化为简笔画式推理,并微调基础多模态模型,赋予其初步的简笔画推理能力;其次训练SketchJudge奖励模型,显式评估模型的思考过程,对简笔画式推理给予更高评分;最后在SketchJudge监督下进行草图思维强化学习,进一步泛化该能力。在四个基准上的实验表明,SketchThinker-R1实现推理令牌成本降低超过64%,且不影响最终答案准确性。定性分析还显示,简笔画式推理在求解过程中更聚焦关键线索。

原文摘要 · Abstract (English)

Despite the empirical success of extensive, step-by-step reasoning in large multimodal models, long reasoning processes inevitably incur substantial computational overhead, i.e., in terms of higher token costs and increased response time, which undermines inference efficiency. In contrast, humans often employ sketch-style reasoning: a concise, goal-directed cognitive process that prioritizes salient information and enables efficient problem-solving. Inspired by this cognitive efficiency, we propose SketchThinker-R1, which incentivizes sketch-style reasoning ability in large multimodal models. Our method consists of three primary stages. In the Sketch-Mode Cold Start stage, we convert standard long reasoning process into sketch-style reasoning and finetune base multimodal model, instilling initial sketch-style reasoning capability. Next, we train SketchJudge Reward Model, which explicitly evaluates thinking process of model and assigns higher scores to sketch-style reasoning. Finally, we conduct Sketch-Thinking Reinforcement Learning under supervision of SketchJudge to further generalize sketch-style reasoning ability. Experimental evaluation on four benchmarks reveals that our SketchThinker-R1 achieves over 64% reduction in reasoning token cost without compromising final answer accuracy. Qualitative analysis further shows that sketch-style reasoning focuses more on key cues during problem solving.

多模态推理效率优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。