提升多模态模型长链反思推理能力,构建新基准并提出自适应训练方法。
MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization
- 设计数据生成引擎构建MM-HELIX基准,包含42类复杂任务
- 在新基准上实现18.6%准确率提升,通用数学逻辑任务增5.7%
- 提出动态融合离线监督与在线优化的AHPO策略,解决稀疏奖励问题
当前多模态大语言模型虽在数学与逻辑推理任务中表现良好,但对解决复杂现实问题所必需的长链反思推理能力仍研究不足。本文通过精心设计的数据合成引擎,构建了包含1,260个样本、42类挑战性合成任务的多模态基准MM-HELIX,用于评估该能力。实证结果表明,现有MLLMs在此类任务上存在显著性能短板。为此,我们生成后训练数据,探索有效学习范式。首先构建Step-Elicited Response Generation流程,创建包含10万条高质量反思推理轨迹的MM-HELIX-100K数据集,用于指令微调阶段。针对强化学习因奖励信号稀疏及监督微调后灾难性遗忘而失效的问题,提出自适应混合策略优化(AHPO),动态融合离线监督与在线优化至单一阶段。该策略使模型在奖励稀疏时学习专家数据,熟练后自主探索。应用于Qwen2.5-VL-7B基线模型,在MM-HELIX基准上实现+18.6%准确率提升,并在通用数学与逻辑任务上平均增益+5.7%。本工作证明多模态模型的反思推理能力可被有效学习与泛化,为发展更强大多模态模型铺平道路。
原文摘要 · Abstract (English)
While current Multimodal Large Language Models (MLLMs) have demonstrated proficiency in reasoning tasks such as mathematics and logic, their capacity for long-chain reflective reasoning, a prerequisite for solving complex real-world problems, remains largely underexplored. In this work, we first conduct an extensive empirical investigation to evaluate this capability. Leveraging a carefully designed data synthesis engine, we construct MM-HELIX, a multimodal benchmark consisting 1,260 samples of 42 challenging synthetic tasks that require iterative thinking and backtracking. Empirical results on this benchmark reveal that existing MLLMs exhibit significant performance deficits in long-chain reflective reasoning. To address this limitation, we generate post-training data and further explore learning paradigms for exploiting such data. We first develop the Step-Elicited Response Generation pipeline to create MM-HELIX-100K, a large-scale dataset of 100k high-quality, reflective reasoning traces for instruction-tuning stage. Given that standard Reinforcement Learning fails on complex tasks due to sparse reward signals and catastrophic forgetting after Supervised Fine-Tuning, we propose Adaptive Hybrid Policy Optimization (AHPO), a novel training strategy that dynamically unifies offline supervision and online optimization into a single stage. This strategy enables the model to learn from expert data when rewards are sparse and conduct independent exploration once proficient. When applied to the Qwen2.5-VL-7B baseline, our method achieves a +18.6\% accuracy improvement on MM-HELIX benchmark and demonstrates strong generalization with a +5.7\% average performance gain on general mathematic and logic tasks. Our work demonstrate that reflective reasoning in MLLMs can be effectively learned and generalized, paving the way for developing more capable MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。