提出新方法提升NLP公平性,兼顾模型性能与可解释性。
Advancing Fairness in Natural Language Processing: From Traditional Methods to Explainability
- 设计新算法降低多分类任务中的偏见,适用于高风险场景。
- 发现数据集规模影响歧视性偏见,传统公平指标存在局限。
- 开发可解释方法识别并排序Transformer中的关键概念,适合研究偏见根源。
自然语言处理(NLP)正面临公平性整合的关键阶段。本博士论文致力于构建更公平、透明的NLP系统,强调其不仅是技术问题,更是道德责任。首先,提出一种新型算法,在高风险NLP应用中有效缓解多分类器偏见,同时保持预测准确率。其次,对Bios数据集的分析揭示了数据规模对歧视性偏见的影响,并指出标准公平性度量的不足。由此推动可解释AI研究,旨在突破传统度量局限。为此,提出COCKATIEL——一种模型无关的可解释方法,能识别并排序Transformer模型中的概念,在情感分析任务中优于现有方法。最后,提出TaCo,一种中和Transformer嵌入偏见的新方法。本研究通过融合公平性与可解释性,挑战现有NLP范式,为构建更公正、负责任的人工智能提供可行方案。
原文摘要 · Abstract (English)
The burgeoning field of Natural Language Processing (NLP) stands at a critical juncture where the integration of fairness within its frameworks has become an imperative. This PhD thesis addresses the need for equity and transparency in NLP systems, recognizing that fairness in NLP is not merely a technical challenge but a moral and ethical necessity, requiring a rigorous examination of how these technologies interact with and impact diverse human populations. Through this lens, this thesis undertakes a thorough investigation into the development of equitable NLP methodologies and the evaluation of biases that prevail in current systems. First, it introduces an innovative algorithm to mitigate biases in multi-class classifiers, tailored for high-risk NLP applications, surpassing traditional methods in both bias mitigation and prediction accuracy. Then, an analysis of the Bios dataset reveals the impact of dataset size on discriminatory biases and the limitations of standard fairness metrics. This awareness has led to explorations in the field of explainable AI, aiming for a more complete understanding of biases where traditional metrics are limited. Consequently, the thesis presents COCKATIEL, a model-agnostic explainability method that identifies and ranks concepts in Transformer models, outperforming previous approaches in sentiment analysis tasks. Finally, the thesis contributes to bridging the gap between fairness and explainability by introducing TaCo, a novel method to neutralize bias in Transformer model embeddings. In conclusion, this thesis constitutes a significant interdisciplinary endeavor that intertwines explicability and fairness to challenge and reshape current NLP paradigms. The methodologies and critiques presented contribute to the ongoing discourse on fairness in machine learning, offering actionable solutions for more equitable and responsible AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。