arXiv:2604.10228cs.AI2026-04被引 2

让多模态模型学会自我检查和修正,提升复杂视觉推理的准确性与可靠性。

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

论文配图:SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
图 1 · 摘自论文原文
  • 引入自我验证与修正机制,构建三阶段训练框架增强推理链条
  • 在多个基准上显著提升推理准确率,且对未见任务泛化能力强
  • 适合追求高可靠性的多模态系统研发者,尤其关注模型可解释性

当前多模态模型常因浅层推理导致错误,源于思维过程不完整或不一致。为此,我们提出自验证与自修正(SVSR)框架,将自我验证与自我修正显式融入推理流程,显著提升复杂视觉理解与多模态推理任务中的鲁棒性与可靠性。SVSR基于新颖的三阶段训练范式:首先,通过精炼预训练视觉语言模型的推理轨迹,构建高质量统一偏好数据集,融合正向与逆向推理以嵌入自我反思信号;其次,在该数据集上进行冷启动监督微调,学习结构化多步推理行为;第三,采用半在线直接偏好优化(Semi-online DPO),持续用强教师视觉语言模型筛选后的高质量模型生成推理轨迹,扩充训练语料。该流程使模型学会生成、激发并优化自我验证与修正能力。大量实验表明,SVSR在多个基准上提升推理准确率,并实现对未见任务和问题类型的更强泛化能力。值得注意的是,经过显式自我反思推理训练后,模型即使未提供显式推理轨迹,也展现出更强的隐式推理能力,超越多个强基线。这些结果凸显了SVSR在构建更可靠、具内省能力、认知对齐的多模态系统方面的潜力。

原文摘要 · Abstract (English)

Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this limitation, we propose Self-Verification and Self-Rectification (SVSR), a unified framework that explicitly integrates self-verification and self-rectification into the model's reasoning pipeline, substantially improving robustness and reliability in complex visual understanding and multimodal reasoning tasks. SVSR is built on a novel three-stage training paradigm. First, we construct a high-quality unified preference dataset by refining reasoning traces from pre-trained vision-language models, incorporating both forward and backward reasoning to embed self-reflective signals. Second, we perform cold-start supervised fine-tuning on this dataset to learn structured, multi-step reasoning behaviors. Third, we apply a Semi-online Direct Preference Optimization (Semi-online DPO) process, continuously augmenting the training corpus with high-quality, model-generated reasoning traces filtered by a powerful teacher VLM. This pipeline enables the model to learn, elicit, and refine its ability to self-verify and self-rectify. Extensive experiments across diverse benchmarks demonstrate that SVSR improves reasoning accuracy and enables stronger generalization to unseen tasks and question types. Notably, once trained with explicit self-reflective reasoning, the model also exhibits improved implicit reasoning ability, outperforming strong baselines even when no explicit reasoning traces are provided. These results highlight the potential of SVSR for building more dependable, introspective, and cognitively aligned multimodal systems.

多模态推理自我修正视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。