arXiv:2605.24794cs.CVcs.CL2026-05

通过对抗自博弈提升视觉语言模型的多模态推理能力。

DUEL: Adversarial Self-Play for Multimodal Reasoning

论文配图:DUEL: Adversarial Self-Play for Multimodal Reasoning
图 1 · 摘自论文原文
  • 两策略对抗生成与验证,自动生成监督信号。
  • 在多个基准上显著提升视觉推理准确率,无需人工标注。
  • 适合研究多模态模型训练与无监督学习的学者。

强化学习(RL)已成为提升视觉-语言模型(VLMs)推理能力的有效范式。然而,基于RL的优化通常依赖于成本高昂的高质量标注数据,难以扩展。现有无监督方法因视觉定位弱和缺乏可靠验证信号,易产生偏差解。我们提出自演化后训练框架DUEL,两个从同一预训练VLM初始化的策略进行对抗交互:挑战者生成一个图像相关的真命题及其最小扰动的难负例,求解者则对两者进行图像验证,促使在近邻语义下实现细粒度视觉区分。为稳定优化,引入长度归一化的对数似然奖励,保留二元结果之外的信息性优化信号,在稀疏反馈下提升学习稳定性。实验表明,DUEL在无需额外人工标注、外部奖励模型或图像编辑工具的情况下,持续提升视觉推理能力和鲁棒性区分能力。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-based optimization typically depends on costly high-quality annotations that are difficult to scale. Existing unsupervised alternatives may drift toward biased solutions due to weak visual grounding and the lack of reliable verification signals. We propose a self-evolving post-training framework, DUEL, where supervision emerges from adversarial interactions between two policies initialized from the same pretrained VLM. A Challenger generates an image-grounded true claim together with a minimally perturbed hard-negative counterpart, while a Solver verifies both claims against the image, encouraging fine-grained visual discrimination under near-neighbor semantics. To stabilize optimization, we introduce a length-normalized log-likelihood reward that preserves informative optimization signals beyond binary outcome supervision and improves learning stability under sparse feedback. Experiments show that DUEL consistently improves visual reasoning and robust discrimination without additional human annotations, external reward models, or image editing tools.

多模态推理对抗训练自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。