arXiv:2506.19028cs.CLcs.AI2025-06被引 5

提出细粒度语义对比方法,更准确检测大模型在长文本中的隐性偏见

Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective

  • 将生成内容拆解为语义命题,通过蕴含判断分析跨群体差异
  • 在性别、种族、年龄三类数据上验证,比传统方法更稳定识别细微偏见
  • 适合关注模型公平性评估的研究者与开发者使用

大型语言模型常生成带有内在偏见的回应,影响其在真实场景中的可靠性。现有评估方法往往忽略长文本中的偏见以及模型输出的固有变异性。为此,我们提出FiSCo(细粒度语义对比)——一种新型统计框架,通过检测不同人口群体在长文本回应中的细微语义差异,评估模型的群体公平性。不同于以往聚焦情感或词级比较的工作,FiSCo在命题层面操作,利用蕴含判断评估语义一致性。我们将模型输出分解为语义独立的命题,应用统计假设检验比较组间与组内相似性,实现对细微偏见的稳健检测。我们形式化了一种新的群体反事实公平定义,并在涵盖性别、种族和年龄的合成与人工标注数据集上验证了FiSCo的有效性。实验表明,该方法更可靠地识别出细微偏见,同时降低随机性带来的干扰,优于多种现有评估指标。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often generate responses with inherent biases, undermining their reliability in real-world applications. Existing evaluation methods often overlook biases in long-form responses and the intrinsic variability of LLM outputs. To address these challenges, we propose FiSCo (Fine-grained Semantic Comparison), a novel statistical framework to evaluate group-level fairness in LLMs by detecting subtle semantic differences in long-form responses across demographic groups. Unlike prior work focusing on sentiment or token-level comparisons, FiSCo goes beyond surface-level analysis by operating at the claim level, leveraging entailment checks to assess the consistency of meaning across responses. We decompose model outputs into semantically distinct claims and apply statistical hypothesis testing to compare inter- and intra-group similarities, enabling robust detection of subtle biases. We formalize a new group counterfactual fairness definition and validate FiSCo on both synthetic and human-annotated datasets spanning gender, race, and age. Experiments show that FiSCo more reliably identifies nuanced biases while reducing the impact of stochastic LLM variability, outperforming various evaluation metrics.

模型公平性语义分析偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。