为大模型提供公平性与鲁棒性形式化验证框架,保障性别公平与毒性检测可靠性
Language Models That Walk the Talk: A Framework for Formal Fairness Certificates
- 在嵌入空间中形式化建模模型鲁棒性,支持对语义扰动的严格验证
- 可证明在性别相关词汇替换下输出保持公平,且毒性输入始终被正确识别
- 适用于高风险场景如性别偏见缓解和内容安全审核,适合伦理AI研究者使用
随着大型语言模型在高风险应用中的普及,确保其鲁棒性和公平性至关重要。尽管取得成功,大模型仍易受对抗攻击影响,例如通过同义词替换等微小扰动即可改变预测结果,在性别偏见缓解和毒性检测等关键领域构成风险。尽管形式化验证已用于神经网络,但其在大语言模型中的应用仍有限。本文提出一个全面的验证框架,用于认证基于Transformer的语言模型的鲁棒性,重点保障性别公平性及不同性别相关术语下的输出一致性。进一步将该方法扩展至毒性检测,提供形式化保证:经对抗操纵的有毒输入能被一致识别并适当屏蔽,从而确保内容审核系统的可靠性。通过在嵌入空间中形式化鲁棒性,本工作增强了语言模型在伦理人工智能部署和内容管理中的可信度。
原文摘要 · Abstract (English)
As large language models become integral to high-stakes applications, ensuring their robustness and fairness is critical. Despite their success, large language models remain vulnerable to adversarial attacks, where small perturbations, such as synonym substitutions, can alter model predictions, posing risks in fairness-critical areas, such as gender bias mitigation, and safety-critical areas, such as toxicity detection. While formal verification has been explored for neural networks, its application to large language models remains limited. This work presents a holistic verification framework to certify the robustness of transformer-based language models, with a focus on ensuring gender fairness and consistent outputs across different gender-related terms. Furthermore, we extend this methodology to toxicity detection, offering formal guarantees that adversarially manipulated toxic inputs are consistently detected and appropriately censored, thereby ensuring the reliability of moderation systems. By formalizing robustness within the embedding space, this work strengthens the reliability of language models in ethical AI deployment and content moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。