arXiv:2607.11228cs.CYcs.AI2026-07

用动态演化测试暴露大模型深层社会偏见,比传统方法更深入。

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

论文配图:DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
图 1 · 摘自论文原文
  • 设计智能代理,通过生成-进化-探测循环自适应生成测试样例。
  • 在5个主流LVLM上验证,揭示了传统静态测试无法发现的深层偏见。
  • 适合关注模型安全、偏见检测的研究者与工程师使用。

大型视觉语言模型(LVLMs)虽能力突出,但仍易受社会偏见影响。现有评估方法多依赖静态数据集,仅能提供表面化评测,无法动态反映模型真实脆弱性边界。本文提出DeepBias,一种自适应深度探测社会偏见的框架,包含精心设计的智能代理。其采用动态‘生成-进化-探测’循环:首先,生成式ProposerAgent基于目标LVLM的反馈,通过直接偏好优化(DPO)迭代更新,探索特定模型的失效模式;其次,自主的技能驱动型DiggerAgent在多轮探测中对测试样本进行重写,从预设技能库中选择深化策略,并根据模型先前响应动态调整。我们构建了基准DeepBiasBench,以五种不同先进LVLM为锚点,捕捉跨架构的共性漏洞。全面实验表明,该框架有效暴露深层偏见,为深度偏见评估提供了挑战性基准,建立了LVLM安全评估的演化范式。

原文摘要 · Abstract (English)

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rewriting strategies. At each turn, this process is conditioned on the model's previous response, enabling progressively deeper biases to be exposed. Furthermore, we build a benchmark named DeepBiasBench using our framework. By employing an ensemble of five diverse state-of-the-art LVLMs as anchors, the benchmark captures vulnerabilities shared across architectures. Comprehensive experiments demonstrate the effectiveness of our framework and show that DeepBias provides a challenging benchmark for in-depth bias evaluation, establishing an evolutionary paradigm for LVLM safety assessment.

模型安全偏见检测LVLM动态测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。