arXiv:2512.00556cs.SEcs.CL2025-12被引 1

用变换规则检测大模型隐性偏见,并通过生成对抗样本实现精准纠错。

Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations

  • 基于变异测试原理设计六种输入变换规则,生成语义等价但具挑战性的偏见样本。
  • 在6个主流大模型上检测到比现有工具多14%的隐藏偏见,且纠偏后安全回应率提升至88.9%。
  • 适合关注大模型公平性、需自动化评估与优化偏见的研究者与开发者使用。

大型语言模型(LLM)广泛应用加剧了其输出中潜藏社会偏见的担忧。现有防护机制在面对间接或情境复杂的偏见提示时常失效。为此,我们提出一个统一框架,用于系统化评估与针对性缓解偏见。该方法引入六种新型元变换关系(Metamorphic Relations, MRs),依据元测试原则,将直接偏见诱导输入转化为语义等价但具有对抗性挑战的变体。这些变换可自动暴露模型隐含偏见:当模型在不同MR变体间响应不一致或不公平时,偏见即被揭示。进一步证明,相同MR可用于生成多样偏见诱导样本进行微调,实现测试与缓解的闭环。在六个先进LLM(涵盖开源与专有模型)及来自8,978项的BiasAsker基准中385个代表性问题(覆盖七个受保护群体)上,本方法比现有工具最多发现14%更多隐藏偏见。此外,结合原始样本与MR生成样本微调后,各模型的安全回应率从54.7%显著提升至超过88.9%。结果表明,元变换关系是提升对话式AI公平性的实用机制。

原文摘要 · Abstract (English)

The widespread deployment of Large Language Models (LLMs) has intensified concerns about subtle social biases embedded in their outputs. Existing guardrails often fail when faced with indirect or contextually complex bias-inducing prompts. To address these limitations, we propose a unified framework for both systematic bias evaluation and targeted mitigation. Our approach introduces six novel Metamorphic Relations (MRs) that, based on metamorphic testing principles, transform direct bias-inducing inputs into semantically equivalent yet adversarially challenging variants. These transformations enable an automated method for exposing hidden model biases: when an LLM responds inconsistently or unfairly across MR-generated variants, the underlying bias becomes detectable. We further show that the same MRs can be used to generate diverse bias-inducing samples for fine-tuning, directly linking the testing process to mitigation. Using six state-of-the-art LLMs - spanning open-source and proprietary models - and a representative subset of 385 questions from the 8,978-item BiasAsker benchmark covering seven protected groups, our MRs reveal up to 14% more hidden biases compared to existing tools. Moreover, fine-tuning with both original and MR-mutated samples significantly enhances bias resiliency, increasing safe response rates from 54.7% to over 88.9% across models. These results highlight metamorphic relations as a practical mechanism for improving fairness in conversational AI.

大模型偏见元变换关系公平性评测自动化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。