让AI推理更靠谱:通过自我评分提升视觉问答的连贯性与准确性
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- 用模型自身输出生成五个自参照信号,自动评估推理过程可靠性
- 在多个视觉基准上,70亿参数模型平均准确率达81.4%,领先开源同类模型
- 无需人工标注,适合希望提升多模态推理可信度的研究者和开发者
多模态大模型常生成流畅但不可靠的推理,其步骤间连贯性弱且视觉对齐不足,主因是现有对齐方法仅监督最终答案,忽略中间推理过程的可靠性。我们提出SR-MCR,一种轻量级、无标签的框架,通过直接从模型输出中提取内在过程信号实现推理对齐。五种自参考线索——语义一致性、词汇忠实度、非冗余性、视觉对齐性与步骤一致性——被整合为归一化的、可靠性加权的奖励,提供细粒度的过程级指导。采用无评论家的GRPO目标,并引入置信度感知降温机制,进一步稳定训练并抑制平庸或过度自信的生成。基于Qwen2.5-VL,SR-MCR在广泛视觉基准上同时提升答案准确率与推理连贯性;在同规模开源模型中,SR-MCR-7B达到81.4%的平均准确率,性能领先。消融实验验证了各奖励项及降温模块的独立贡献。
原文摘要 · Abstract (English)
Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approaches supervise only the final answer while ignoring the reliability of the intermediate reasoning process. We introduce SR-MCR, a lightweight and label-free framework that aligns reasoning by exploiting intrinsic process signals derived directly from model outputs. Five self-referential cues -- semantic alignment, lexical fidelity, non-redundancy, visual grounding, and step consistency -- are integrated into a normalized, reliability-weighted reward that provides fine-grained process-level guidance. A critic-free GRPO objective, enhanced with a confidence-aware cooling mechanism, further stabilizes training and suppresses trivial or overly confident generations. Built on Qwen2.5-VL, SR-MCR improves both answer accuracy and reasoning coherence across a broad set of visual benchmarks; among open-source models of comparable size, SR-MCR-7B achieves state-of-the-art performance with an average accuracy of 81.4%. Ablation studies confirm the independent contributions of each reward term and the cooling module.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。