arXiv:2512.05530cs.AI2025-12中稿 · ICML

让多模态大模型像人一样思考:理解-反思-修正,提升推理可靠性。

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models

  • 引入多理由增强与判别机制,构建统一数据基础。
  • 分阶段训练实现逻辑纠错,准确率显著提升。
  • 适合需要高可靠推理的多模态应用开发者。

近期,多模态大语言模型(MLLMs)被广泛应用于推理任务,但存在多理由语义建模能力有限、逻辑鲁棒性不足及易受误导线索影响的问题。为此,我们提出一种多理由集成判别推理框架(MIND),旨在赋予MLLMs类人的“理解→重思→修正”认知能力,实现从被动模仿到主动判别推理的范式演进。具体而言,我们设计了理由增强与判别(RAD)范式,提供统一可扩展的数据基础;并提出渐进式两阶段修正学习(P2CL)策略,第一阶段强化多理由正向学习,第二阶段实现主动逻辑判别与修正。此外,为缓解多理由语义空间中的表示纠缠,我们提出了多理由对比对齐(MCA)优化策略。大量实验表明,MIND在多个公开数据集上达到当前最优性能。数据与代码已开源于https://github.com/YuChuang1205/MIND。

原文摘要 · Abstract (English)

Recently, multimodal large language models (MLLMs) have been widely applied to reasoning tasks. However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and susceptibility to misleading cues. Therefore, we propose a Multi-rationale INtegrated Discriminative (MIND) reasoning framework, which is designed to endow MLLMs with human-like cognitive abilities of "Understand -> Rethink -> Correct", and achieves a paradigm evolution from passive imitation-based reasoning to active discriminative reasoning. Specifically, we introduce a Rationale Augmentation and Discrimination (RAD) paradigm, which provides a unified and extensible data foundation. Meanwhile, we design a Progressive Two-stage Correction Learning (P2CL) strategy. The first phase enhances multi-rationale positive learning, while the second phase enables active logic discrimination and correction. In addition, to mitigate representation entanglement in the multi-rationale semantic space, we propose a Multi-rationale Contrastive Alignment (MCA) optimization strategy. Extensive experiments show that our MIND achieves SOTA performance on multiple public datasets. Our data and code are available at https://github.com/YuChuang1205/MIND

多模态推理逻辑纠错大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。