分析四大闭源AI内容审核模型的公平性与鲁棒性缺陷
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

- 对比四个闭源审核模型在不同群体间的分类偏差
- 发现模型对微小输入变化敏感,存在一致性问题
- 适合关注AI伦理与安全系统的研究人员参考
AI安全审核(ASM)分类器旨在社交平台内容审核中充当护栏,防止大语言模型被不安全输入微调。由于可能造成差异性影响,必须确保这些模型:(1) 不对少数群体用户的内容产生更高误判率;(2) 在相似输入下行为保持稳定一致。本文针对四个广泛使用的闭源ASM分类器——OpenAI Moderation API、Perspective API、Google Cloud Natural Language(GCNL)API和Clarifai API——进行公平性与鲁棒性评估。采用性别、种族等维度的公平性指标如群体平等性和条件统计平等性,对比其表现与公平性基线模型。同时通过测试小规模自然扰动下的响应变化,分析模型鲁棒性。结果揭示了显著的公平性缺口与鲁棒性不足,表明未来版本亟需改进。
原文摘要 · Abstract (English)
AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for disparate impact, it is crucial to ensure that these classifiers: (1) do not unfairly classify content belonging to users from minority groups as unsafe compared to those from majority groups and (2) that their behavior remains robust and consistent across similar inputs. In this work, we thus examine the fairness and robustness of four widely-used, closed-source ASM classifiers: OpenAI Moderation API, Perspective API, Google Cloud Natural Language (GCNL) API, and Clarifai API. We assess fairness using metrics such as demographic parity and conditional statistical parity, comparing their performance against ASM models and a fair-only baseline. Additionally, we analyze robustness by testing the classifiers' sensitivity to small and natural input perturbations. Our findings reveal potential fairness and robustness gaps, highlighting the need to mitigate these issues in future versions of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。