系统梳理大模型偏见的来源、评估与缓解方法,助力公平AI发展。
Bias in Large Language Models: Origin, Evaluation, and Mitigation
- 区分内在与外在偏见,分析其在各类NLP任务中的表现
- 提出数据、模型、输出三层次评估框架,支持精准检测
- 归纳预训练、训练中、后处理三类缓解策略,适合研究者参考
大型语言模型(LLMs)虽推动自然语言处理进步,但其易受偏见影响带来重大挑战。本文全面综述了LLM中偏见的成因、评估与缓解策略。将偏见分为内在与外在两类,分析其在多种NLP任务中的具体表现。批判性评估了数据级、模型级和输出级的偏见检测方法,为研究人员提供实用工具箱。进一步探讨了预模型、模型内与后模型三类缓解技术,阐明其效果与局限性。讨论了偏见模型在医疗与司法等现实应用中的伦理与法律风险。通过整合现有知识,本综述促进公平、负责任的AI发展,为相关研究者与实践者提供系统性参考。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This comprehensive review examines the landscape of bias in LLMs, from its origins to current mitigation strategies. We categorize biases as intrinsic and extrinsic, analyzing their manifestations in various NLP tasks. The review critically assesses a range of bias evaluation methods, including data-level, model-level, and output-level approaches, providing researchers with a robust toolkit for bias detection. We further explore mitigation strategies, categorizing them into pre-model, intra-model, and post-model techniques, highlighting their effectiveness and limitations. Ethical and legal implications of biased LLMs are discussed, emphasizing potential harms in real-world applications such as healthcare and criminal justice. By synthesizing current knowledge on bias in LLMs, this review contributes to the ongoing effort to develop fair and responsible AI systems. Our work serves as a comprehensive resource for researchers and practitioners working towards understanding, evaluating, and mitigating bias in LLMs, fostering the development of more equitable AI technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。