arXiv:2504.07982cs.CLcs.AI2025-04被引 10

用变异测试发现大模型在敏感属性交叉下的偏见问题

Metamorphic Testing for Fairness Evaluation in Large Language Models: Identifying Intersectional Bias in LLaMA and GPT

  • 设计公平性导向的变异关系,生成对比测试用例
  • 在LLaMA和GPT中检测出语气与情感方面的偏见模式
  • 适合关注模型公平性与安全部署的研究者使用

大型语言模型(LLMs)在自然语言处理中取得显著进展,但仍存在公平性问题,常反映训练数据中的固有偏见。这些偏见在医疗、金融、法律等敏感领域部署时带来风险。本文提出一种变异测试方法,系统识别LLMs中的公平性缺陷。定义并应用一系列面向公平性的变异关系(MRs),在多样化人口属性输入下评估LLaMA和GPT模型。通过生成源测试用例与后续测试用例,并分析模型响应中的公平性违规行为,结果表明该方法有效暴露了语气与情感相关的偏见模式,尤其揭示了特定敏感属性交叉组合频繁出现的公平性故障。本研究提升了大模型公平性测试水平,为检测与缓解偏见提供了结构化方法,增强了模型在公平性敏感场景下的鲁棒性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant strides in Natural Language Processing but remain vulnerable to fairness-related issues, often reflecting biases inherent in their training data. These biases pose risks, particularly when LLMs are deployed in sensitive areas such as healthcare, finance, and law. This paper introduces a metamorphic testing approach to systematically identify fairness bugs in LLMs. We define and apply a set of fairness-oriented metamorphic relations (MRs) to assess the LLaMA and GPT model, a state-of-the-art LLM, across diverse demographic inputs. Our methodology includes generating source and follow-up test cases for each MR and analyzing model responses for fairness violations. The results demonstrate the effectiveness of MT in exposing bias patterns, especially in relation to tone and sentiment, and highlight specific intersections of sensitive attributes that frequently reveal fairness faults. This research improves fairness testing in LLMs, providing a structured approach to detect and mitigate biases and improve model robustness in fairness-sensitive applications.

模型公平性变异测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。