arXiv:2512.17532cs.CVcs.AI2025-12AAAI被引 8

让AI模型看清模糊图像,主动分析画质退化原因。

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

  • 用结构化推理链显式分析图像退化的类型和程度
  • 在真实退化数据集上表现超越现有模型,多强度攻击下仍稳定
  • 适合需要高鲁棒性的视觉理解场景,如自动驾驶、医疗影像

多模态大语言模型在极端真实世界视觉退化下难以保持可靠性能,限制了其实际应用。现有鲁棒多模态模型主要依赖隐式训练/适配,仅关注视觉编码器泛化,存在可解释性差、优化孤立的问题。为此,我们提出Robust-R1,一种通过结构化推理链显式建模视觉退化的新型框架。该方法包含:(i) 监督微调构建退化感知推理基础,(ii) 奖励驱动对齐以准确感知退化参数,(iii) 根据退化强度动态调整推理深度。为支持该方法,我们构建了一个包含11,000条样本的专用数据集,涵盖四个关键现实视觉处理阶段的逼真退化,每条样本均标注了连接退化参数、感知影响、原始语义推理链与结论的结构化链条。全面评估显示,Robust-R1在真实退化基准R-Bench上优于所有通用及鲁棒基线,在MMMB、MMStar和RealWorldQA上的多强度对抗退化测试中也保持卓越抗退化能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-the-art robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA.

视觉理解鲁棒性多模态推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。