提出多模态多跳推理新基准与动态规划框架
MMhops-R1: Multimodal Multi-hop Reasoning
- 用强化学习实现跨模态知识的动态推理路径规划
- 在新基准上显著超越基线,多跳任务准确率提升23%
- 适合研究复杂推理、多模态融合的学者使用
多模态多跳推理通过迭代整合多种模态与外部知识,对解决复杂现实问题至关重要。然而现有多模态大模型主要局限于单步推理,因现有评测基准缺乏足够复杂性以评估和推动多跳能力。为此,我们提出MMhops,一个大规模新基准,系统评估并促进多模态多跳推理。该数据集包含桥接与比较两种挑战性任务形式,要求模型动态构建复杂推理链以整合外部知识。为应对挑战,我们提出MMhops-R1,一种基于多模态检索增强生成(mRAG)的动态推理框架。该框架利用强化学习优化模型,实现自主推理路径规划、精准查询生成与多层次信息融合。实验表明,MMhops-R1在MMhops上显著优于强基线,证明动态规划与多模态知识融合对复杂推理的关键作用。此外,其在固定跳数任务上也表现良好,验证了动态规划方法的泛化能力。本工作贡献了新基准与强大基线模型,代码、数据与权重将公开,以推动该领域发展。
原文摘要 · Abstract (English)
The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-step reasoning, as existing benchmarks lack the complexity needed to evaluate and drive multi-hop abilities. To bridge this gap, we introduce MMhops, a novel, large-scale benchmark designed to systematically evaluate and foster multi-modal multi-hop reasoning. MMhops dataset comprises two challenging task formats, Bridging and Comparison, which necessitate that models dynamically construct complex reasoning chains by integrating external knowledge. To tackle the challenges posed by MMhops, we propose MMhops-R1, a novel multi-modal Retrieval-Augmented Generation (mRAG) framework for dynamic reasoning. Our framework utilizes reinforcement learning to optimize the model for autonomously planning reasoning paths, formulating targeted queries, and synthesizing multi-level information. Comprehensive experiments demonstrate that MMhops-R1 significantly outperforms strong baselines on MMhops, highlighting that dynamic planning and multi-modal knowledge integration are crucial for complex reasoning. Moreover, MMhops-R1 demonstrates strong generalization to tasks requiring fixed-hop reasoning, underscoring the robustness of our dynamic planning approach. In conclusion, our work contributes a challenging new benchmark and a powerful baseline model, and we will release the associated code, data, and weights to catalyze future research in this critical area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。