arXiv:2504.00310cs.CLcs.AI2025-04被引 7

用知识图谱训练大模型,有效降低偏见输出。

Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training

  • 引入真实世界知识图谱增强模型训练
  • 在多个数据集上显著降低偏见指标
  • 适合关注AI伦理与公平性的研究者

大型语言模型虽在自然语言处理中表现卓越,但常继承并放大训练数据中的偏见,引发伦理与公平问题。本文提出知识图谱增强训练(KGAT)方法,利用真实世界领域知识图谱中的结构化知识,提升模型理解能力并减少偏见输出。采用Gender Shades、Bias in Bios和FairFace等公开数据集进行评估,以人口均等性和机会平等性为指标。通过针对性修正策略,显著降低偏见输出,改善各项偏见指标。该框架结合真实数据与知识图谱,具备可扩展性与有效性,为敏感场景下的负责任部署提供支持。

原文摘要 · Abstract (English)

Large language models have revolutionized natural language processing with their surprising capability to understand and generate human-like text. However, many of these models inherit and further amplify the biases present in their training data, raising ethical and fairness concerns. The detection and mitigation of such biases are vital to ensuring that LLMs act responsibly and equitably across diverse domains. This work investigates Knowledge Graph-Augmented Training (KGAT) as a novel method to mitigate bias in LLM. Using structured domain-specific knowledge from real-world knowledge graphs, we improve the understanding of the model and reduce biased output. Public datasets for bias assessment include Gender Shades, Bias in Bios, and FairFace, while metrics such as demographic parity and equal opportunity facilitate rigorous detection. We also performed targeted mitigation strategies to correct biased associations, leading to a significant drop in biased output and improved bias metrics. Equipped with real-world datasets and knowledge graphs, our framework is both scalable and effective, paving the way toward responsible deployment in sensitive and high-stakes applications.

大模型偏见检测知识图谱公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。