arXiv:2510.21807cs.CVcs.AI2025-10AAAI

通过遮蔽预测激发视觉上下文与常识推理,提升多模态模型泛化能力。

Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs

  • 设计遮蔽预测任务,强制模型结合视觉上下文与常识推理重建被遮内容。
  • 在跨任务和分布外场景中,模型推理准确率显著提升,验证了泛化性增强。
  • 适用于需要复杂视觉理解与常识推理的多模态应用,如智能客服、机器人交互。

近期推理模型的突破显著提升了大语言模型的推理能力,尤其通过可验证奖励的任务训练实现。然而,在真实世界的多模态场景(尤其是视觉-语言任务)中,仍存在显著差距,主要因过度关注单模态语言设置。尽管已有将强化学习从NLP移植到视觉语言模型(VLMs)的努力,但这些方法通常局限于感知任务或仅将图像简化为文本摘要,未能充分挖掘视觉上下文与常识知识,限制了推理能力在多样化多模态环境中的泛化。为此,我们提出一种新型微调任务:基于上下文与常识的遮蔽预测(Masked Prediction via Context and Commonsense),迫使模型通过重构被遮图像中的语义内容来整合视觉上下文与常识推理,从而奠定通用推理的基础。为系统评估模型在通用推理上的表现,我们构建了专用评估基准MPCC Eval,并采用多种微调策略引导推理。其中,我们引入创新训练方法——先验采样强化微调(Reinforcement Fine tuning with Prior Sampling),不仅提升性能,还显著增强模型在分布外(OOD)和跨任务场景下的泛化推理能力。

原文摘要 · Abstract (English)

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real world multimodal scenarios, most notably, vision language tasks, due to a heavy focus on single modal language settings. While efforts to transplant reinforcement learning techniques from NLP to VLMs have emerged, these approaches often remain confined to perception centric tasks or reduce images to textual summaries, failing to fully exploit visual context and commonsense knowledge, ultimately constraining the generalization of reasoning capabilities across diverse multimodal environments. To address this limitation, we introduce a novel fine tuning task, Masked Prediction via Context and Commonsense, which forces models to integrate visual context and commonsense reasoning by reconstructing semantically meaningful content from occluded images, thereby laying the foundation for generalized reasoning. To systematically evaluate the model performance in generalized reasoning, we developed a specialized evaluation benchmark, MPCC Eval, and employed various fine tuning strategies to guide reasoning. Among these, we introduced an innovative training method, Reinforcement Fine tuning with Prior Sampling, which not only enhances model performance but also improves its generalized reasoning capabilities in OOD and cross task scenarios.

多模态常识推理视觉理解强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。