arXiv:2504.04158cs.CV2025-04CVPR被引 51

用视觉大模型指挥多个修复模型,提升自动驾驶在恶劣天气下的感知能力。

JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration

  • 用视觉大模型做控制器,动态调度多个修复专家模型。
  • 在真实数据上实现50%的感知指标提升,优于现有方法。
  • 适合自动驾驶、智能交通系统等需要强鲁棒性的场景。

以视觉为中心的感知系统在真实环境中面临不可预测且耦合的天气退化问题。当前方案常受限于特定退化先验或存在显著领域差距。为实现真实世界条件下鲁棒且自主的运行,我们提出 JarvisIR,一个由视觉语言模型(VLM)驱动的智能代理,利用 VLM 作为控制器来管理多个专家修复模型。为进一步增强系统鲁棒性、减少幻觉并提升在真实恶劣天气下的泛化能力,JarvisIR 采用新颖的两阶段框架:监督微调与人类反馈对齐。针对真实场景中成对数据稀缺的问题,人类反馈对齐使 VLM 能在大规模真实数据上以无监督方式有效微调。为支持训练与评估,我们构建了 CleanBench,一个包含15万条合成指令-响应对和8万条真实指令-响应对的综合性数据集。大量实验表明,JarvisIR 在决策与修复能力上表现卓越,在 CleanBench-Real 上所有感知指标平均提升50%。

原文摘要 · Abstract (English)

Vision-centric perception systems struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world conditions, we propose JarvisIR, a VLM-powered agent that leverages the VLM as a controller to manage multiple expert restoration models. To further enhance system robustness, reduce hallucinations, and improve generalizability in real-world adverse weather, JarvisIR employs a novel two-stage framework consisting of supervised fine-tuning and human feedback alignment. Specifically, to address the lack of paired data in real-world scenarios, the human feedback alignment enables the VLM to be fine-tuned effectively on large-scale real-world data in an unsupervised manner. To support the training and evaluation of JarvisIR, we introduce CleanBench, a comprehensive dataset consisting of high-quality and large-scale instruction-responses pairs, including 150K synthetic entries and 80K real entries. Extensive experiments demonstrate that JarvisIR exhibits superior decision-making and restoration capabilities. Compared with existing methods, it achieves a 50% improvement in the average of all perception metrics on CleanBench-Real. Project page: https://cvpr2025-jarvisir.github.io/.

自动驾驶图像修复视觉大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。