区分事实正确与公平,避免模型看似合理实则有害。
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases
- 用认知偏差理论设计双类测试题:客观事实型与主观公平型。
- 十款前沿模型均存在事实与公平的权衡,表现各异。
- 适合关注模型社会风险与评估标准的开发者与研究者。
近期事件如谷歌Gemini生成有色人种穿着纳粹制服,凸显了AI输出可能在事实层面看似合理却具有社会危害性。当前对AI模型的‘公平性’评估常混淆两个根本不同的维度:事实正确性与规范性公平性。一个模型可能生成事实准确但社会不公平的内容,或虽表面公平却扭曲事实。我们主张明确区分事实与公平的边界,方能实现有意义的公平评估。为此,提出Fact-or-Fair基准,包含两类查询:(i) 与描述性事实判断对齐的客观问题,(ii) 与规范性公平判断对齐的主观问题。所有问题基于19项统计数据,植根于认知心理学中的代表性偏差、归因偏差和内群体-外群体偏差,解释为何模型常在事实与公平间失衡。在十款前沿模型上的实验揭示了不同程度的事实-公平权衡。本研究不仅提供新的理论视角,也推出可公开使用的评测工具包(https://github.com/uclanlp/Fact-or-Fair),推动负责任的模型评估发展。
原文摘要 · Abstract (English)
Recent failures such as Google Gemini generating people of color in Nazi-era uniforms illustrate how AI outputs can be factually plausible yet socially harmful. AI models are increasingly evaluated for "fairness," yet existing benchmarks often conflate two fundamentally different dimensions: factual correctness and normative fairness. A model may generate responses that are factually accurate but socially unfair, or conversely, appear fair while distorting factual reality. We argue that identifying the boundary between fact and fair is essential for meaningful fairness evaluation. We introduce Fact-or-Fair, a benchmark with (i) objective queries aligned with descriptive, fact-based judgments, and (ii) subjective queries aligned with normative, fairness-based judgments. Our queries are constructed from 19 statistics and are grounded in cognitive psychology, drawing on representativeness bias, attribution bias, and ingroup-outgroup bias to explain why models often misalign fact and fairness. Experiments across ten frontier models reveal different levels of fact-fair trade-offs. By reframing fairness evaluation, we provide both a new theoretical lens and a practical benchmark to advance the responsible model assessments. Our test suite is publicly available at https://github.com/uclanlp/Fact-or-Fair.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。