arXiv:2504.14395cs.CVcs.AI2025-04被引 1

Hydra通过动态推理与多模型验证,同时提升视觉语言模型的抗干扰能力和事实准确性。

Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models

  • 采用迭代式行动-批判循环,结合思维链与上下文学习动态优化输出。
  • 在4个VLM、3个幻觉基准上超越当前最佳去幻觉方法,且无需显式对抗防御。
  • 适用于医疗、安防等高风险场景,无需训练即可增强模型可靠性。

为构建可信的视觉语言模型(VLMs),必须解决对抗鲁棒性与幻觉问题,二者均影响军事、医疗等高风险应用中的事实准确性。现有方法多聚焦于对抗防御或幻觉事后修正,缺乏统一策略。本文提出Hydra,一种自适应代理框架,通过迭代推理、结构化批判与跨模型验证,增强插件式VLM的抗干扰能力与内在误差抑制。Hydra采用行动-批判循环,检索并批判视觉信息,利用思维链(CoT)和上下文学习(ICL)技术动态优化输出。相比静态事后修正,Hydra可应对恶意扰动与模型内在错误,兼具鲁棒性与一致性。我们在4个VLM、3个幻觉基准、2种对抗攻击策略及2种防御方法上评估,涵盖清洁与对抗输入。结果表明,Hydra在未显式部署对抗防御的情况下,仍优于插件式VLM与当前最优去幻觉方法,显著提升鲁棒性与事实一致性。Hydra实现了对抗防御与幻觉缓解的统一,提供了一种可扩展、免训练的解决方案,适用于真实世界VLM可靠性提升。

原文摘要 · Abstract (English)

To develop trustworthy Vision-Language Models (VLMs), it is essential to address adversarial robustness and hallucination mitigation, both of which impact factual accuracy in high-stakes applications such as defense and healthcare. Existing methods primarily focus on either adversarial defense or hallucination post-hoc correction, leaving a gap in unified robustness strategies. We introduce \textbf{Hydra}, an adaptive agentic framework that enhances plug-in VLMs through iterative reasoning, structured critiques, and cross-model verification, improving both resilience to adversarial perturbations and intrinsic model errors. Hydra employs an Action-Critique Loop, where it retrieves and critiques visual information, leveraging Chain-of-Thought (CoT) and In-Context Learning (ICL) techniques to refine outputs dynamically. Unlike static post-hoc correction methods, Hydra adapts to both adversarial manipulations and intrinsic model errors, making it robust to malicious perturbations and hallucination-related inaccuracies. We evaluate Hydra on four VLMs, three hallucination benchmarks, two adversarial attack strategies, and two adversarial defense methods, assessing performance on both clean and adversarial inputs. Results show that Hydra surpasses plug-in VLMs and state-of-the-art (SOTA) dehallucination methods, even without explicit adversarial defenses, demonstrating enhanced robustness and factual consistency. By bridging adversarial resistance and hallucination mitigation, Hydra provides a scalable, training-free solution for improving the reliability of VLMs in real-world applications.

视觉语言模型对抗鲁棒性幻觉抑制推理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。