arXiv:2505.07870cs.CLcs.AI2025-05被引 6

用句子多样性优先排序检测大模型偏见,效率提升显著。

Efficient Fairness Testing in Large Language Models: Prioritizing Metamorphic Relations for Bias Detection

  • 基于句子多样性排序元测试关系,优化偏见检测
  • 故障发现率提升22%,首次失败时间缩短15%
  • 无需标注故障,适合大规模模型公平性测试

大型语言模型(LLMs)在各类应用中日益普及,其输出中的公平性与潜在偏见引发广泛关注。本文探讨了在元测试(metamorphic testing)中对元测试关系(MRs)进行优先级排序的策略,以高效检测LLMs中的公平性问题。由于可能的测试用例呈指数级增长,全面测试不切实际;因此,依据检测公平性违规的有效性对MRs进行优先排序至关重要。我们采用基于句子多样性的方法计算并排序MRs,以优化故障检测。实验结果表明,所提出的优先级排序方法相较于随机排序,故障发现率提升22%,首次失败时间缩短15%;相比基于距离的排序,故障发现率提升12%,首次失败时间缩短8%。此外,该方法在有效性上仅比基于故障的排序低5%,但显著降低了故障标注带来的计算成本。这些结果验证了基于多样性的MR优先级排序在提升LLM公平性测试方面的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in various applications, raising critical concerns about fairness and potential biases in their outputs. This paper explores the prioritization of metamorphic relations (MRs) in metamorphic testing as a strategy to efficiently detect fairness issues within LLMs. Given the exponential growth of possible test cases, exhaustive testing is impractical; therefore, prioritizing MRs based on their effectiveness in detecting fairness violations is crucial. We apply a sentence diversity-based approach to compute and rank MRs to optimize fault detection. Experimental results demonstrate that our proposed prioritization approach improves fault detection rates by 22% compared to random prioritization and 12% compared to distance-based prioritization, while reducing the time to the first failure by 15% and 8%, respectively. Furthermore, our approach performs within 5% of fault-based prioritization in effectiveness, while significantly reducing the computational cost associated with fault labeling. These results validate the effectiveness of diversity-based MR prioritization in enhancing fairness testing for LLMs.

公平性测试元测试大模型偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。