让AI看图推理更靠谱,避免错答却有理有据的幻觉。
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
- 构建带结构化伪真实数据的感知对齐层,让推理过程可验证。
- 设计分层奖励融合机制,使视觉推理链忠实于图像证据。
- 在多个评测上显著降低幻觉,适合追求可信AI的开发者。
强化学习虽提升了大语言模型与多模态模型的推理能力,但现有奖励机制仅关注最终答案正确性,容忍推理过程中的视觉幻觉——即模型得出正确答案却误解了视觉信息。为此,本文提出PaLMR框架,不仅对齐结果,更对齐推理过程。该框架包含两部分:感知对齐数据层,构建含结构化伪真实与可验证视觉事实的过程感知数据;过程对齐优化层,设计分层奖励融合机制与过程感知评分函数,鼓励视觉忠实的思维链并提升训练稳定性。在Qwen2.5-VL-7B上的实验表明,该方法显著减少推理幻觉,提升视觉推理保真度,在HallusionBench上达到当前最佳性能,同时保持在MMMU、MathVista和MathVerse上的强表现。结果表明,PaLMR为多模态推理提供了原理清晰且实用的对齐路径,提升了多模态大模型的可靠性与可解释性。
原文摘要 · Abstract (English)
Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where models reach the right answer while misperceiving visual evidence. We address this process-level misalignment with PaLMR, a framework that aligns not only outcomes but also the reasoning process itself. PaLMR comprises two complementary components: a perception-aligned data layer that constructs process-aware reasoning data with structured pseudo-ground-truths and verifiable visual facts, and a process-aligned optimisation layer that constructs a hierarchical reward fusion scheme with a process-aware scoring function to encourage visually faithful chains-of-thought and improve training stability. Experiments on Qwen2.5-VL-7B show that our approach substantially reduces reasoning hallucinations and improves visual reasoning fidelity, achieving state-of-the-art results on HallusionBench while maintaining strong performance on MMMU, MathVista, and MathVerse. These findings indicate that PaLMR offers a principled and practical route to process-aligned multimodal reasoning, advancing the reliability and interpretability of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。