arXiv:2504.12898cs.CLcs.AI2025-04

用信息增益指导因果干预,自动消除指令微调数据偏见

Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models

  • 基于信息增益为0的约束,通过因果干预重写数据分布
  • 在多个任务上显著提升模型泛化能力,验证去偏效果
  • 适合需要高公平性、低偏见的LLM应用开发者

尽管取得进展,现有大语言模型仍可能在推理中利用训练数据中的偏见,导致泛化能力差。由于数据偏见多样且基于上下文学习的去偏方法有限,以往依赖先验知识或上下文学习的去偏方法效果受限。为此,本文提出信息增益引导的因果干预去偏框架(ICD),核心思想是:指令微调数据中的偏见不应提供额外预测答案的信息,即信息增益应为0。基于此,框架采用因果干预方法自动重写数据,平衡其分布以降低信息增益,随后使用标准监督微调训练模型。实验表明,ICD能有效去偏,显著提升模型在不同任务上的泛化性能。

原文摘要 · Abstract (English)

Despite significant progress, recent studies indicate that current large language models (LLMs) may still capture dataset biases and utilize them during inference, leading to the poor generalizability of LLMs. However, due to the diversity of dataset biases and the insufficient nature of bias suppression based on in-context learning, the effectiveness of previous prior knowledge-based debiasing methods and in-context learning based automatic debiasing methods is limited. To address these challenges, we explore the combination of causal mechanisms with information theory and propose an information gain-guided causal intervention debiasing (ICD) framework. To eliminate biases within the instruction-tuning dataset, it is essential to ensure that these biases do not provide any additional information to predict the answers, i.e., the information gain of these biases for predicting the answers needs to be 0. Under this guidance, this framework utilizes a causal intervention-based data rewriting method to automatically and autonomously balance the distribution of instruction-tuning dataset for reducing the information gain. Subsequently, it employs a standard supervised fine-tuning process to train LLMs on the debiased dataset. Experimental results show that ICD can effectively debias LLM to improve its generalizability across different tasks.

去偏因果干预语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。