arXiv:2507.06999cs.CVcs.CL2025-07被引 1

用格式奖励训练多模态模型,推理时自动灵活应变。

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

  • 训练时用规则格式奖励引导深度思考,不需额外标注
  • 测试时切换为直觉式响应,性能超越基线模型
  • 适合追求高效推理的多模态系统开发者

推理对大型语言模型至关重要,尤其在数学问题求解等复杂任务中。然而,多模态推理仍面临模态对齐与训练可扩展性的挑战,现有方法常依赖额外标注或复杂的基于规则的奖励机制。为此,我们提出刻意-直觉推理框架(D2I),在无需额外标注或复杂奖励的情况下提升多模态大模型(MLLM)的理解与推理能力。训练阶段,D2I仅通过基于规则的格式奖励监督刻意推理策略,增强模态对齐;推理阶段则移除显式推理策略,使模型以隐式方式应用所学能力。D2I在域内与域外基准测试中均优于基线模型,证明格式奖励能有效培养可迁移的多模态推理技能,并支持训练时深度推理与测试时响应灵活性的解耦。

原文摘要 · Abstract (English)

Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal reasoning still faces challenges in modality alignment and training scalability, as many existing methods rely on additional annotations or complex rule-based rewards. To address these issues, we propose the Deliberate-to-Intuitive reasoning framework (D2I), which improves the understanding and reasoning abilities of multimodal LLMs (MLLMs) without extra annotations or complex rewards. During training, D2I uses deliberate reasoning strategies supervised only by rule-based format rewards to enhance modality alignment. During inference, it shifts to intuitive reasoning by removing these explicit strategies, allowing the model to implicitly apply the acquired abilities in its responses. D2I outperforms baselines on both in-domain and out-of-domain benchmarks, highlighting the effectiveness of format rewards in fostering transferable multimodal reasoning skills and suggesting the benefit of decoupling training-time reasoning depth from test-time response flexibility.

多模态推理语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。