arXiv:2602.22227cs.LGcs.AI2026-02

用自演化对抗训练提升多模态大模型的视觉鲁棒性

Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models

  • 构建攻击者-防御者自对弈框架,动态生成对抗样本
  • 在多个视觉任务上显著降低幻觉率,提升模型稳健性
  • 适合关注多模态模型可靠性与对抗训练的研究者

尽管多模态大语言模型(MLLMs)能力出众,但在面对复杂视觉场景时仍存在感知脆弱性。这一缺陷源于对有限训练数据集的依赖,而这些数据集因成本高昂难以扩展,限制了模型鲁棒性的提升。我们提出 extbf{AOT-SFT},一个大规模对抗数据集,用于启动 MLLM 的鲁棒性训练。基于此,我们进一步设计 extbf{AOT(Adversarial Opponent Training)},一种自对弈框架,通过让图像编辑攻击者与防御型 MLLM 协同进化,自动生成多样化且动态的图像扰动作为训练课程。该机制迫使防御模型持续适应新挑战,实现性能迭代。大量实验表明,AOT 显著提升了防御模型的感知鲁棒性并减少了幻觉现象,为训练更可靠的 MLLMs 提供了一种可扩展的新范式。

原文摘要 · Abstract (English)

Despite their impressive capabilities, Multimodal Large Language Models (MLLMs) exhibit perceptual fragility when confronted with visually complex scenes. This weakness stems from a reliance on finite training datasets, which are prohibitively expensive to scale and impose a ceiling on model robustness. We introduce \textbf{AOT-SFT}, a large-scale adversarial dataset for bootstrapping MLLM robustness. Building on this, we propose \textbf{AOT (Adversarial Opponent Training)}, a self-play framework that forges MLLM robustness by creating its own training data. Our method orchestrates a co-evolution between an image-editing Attacker and a Defender MLLM, where the Attacker generates a diverse and dynamic curriculum of image manipulations, forcing the Defender to adapt and improve. Extensive experiments demonstrate that AOT enhances the Defender's perceptual robustness and reduces hallucinations, establishing a scalable paradigm for training more reliable MLLMs.

多模态对抗训练鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。