arXiv:2502.18249cs.LGcs.CL2025-02AAAI被引 2

通过迭代反事实增强,有效去除数据偏见并提升模型推理与人工标注的一致性。

Iterative Counterfactual Data Augmentation

  • 采用初始高噪声干预的迭代反事实增强方法
  • 显著降低冗余信息,保留目标信号与标签的互信息
  • 适用于需要可解释性与公平性的文本任务

反事实数据增强(CDA)是一种通过生成具有相反偏见的互补数据集来控制训练数据中信息或偏见的方法。以往工作多依赖手工规则或算法生成,常在增强数据中残留不希望的信息。本文提出迭代反事实增强(ICDA),通过初始高噪声干预,可收敛至噪声显著降低的状态。ICDA生成的数据集使目标信号与对应标签保持高互信息,同时减少虚假信号的信息量。在六个人工标注数据集和两个大语言模型生成数据集上的实验表明,使用增强数据训练的模型在文档推理上更符合人工标注。关键在于:高噪声初始干预能逐步净化数据,实现更可信的模型行为。

原文摘要 · Abstract (English)

Counterfactual data augmentation (CDA) is a method for controlling information or biases in training datasets by generating a complementary dataset with typically opposing biases. Prior work often either relies on hand-crafted rules or algorithmic CDA methods which can leave unwanted information in the augmented dataset. In this work, we show iterative CDA (ICDA) with initial, high-noise interventions can converge to a state with significantly lower noise. Our ICDA procedure produces a dataset where one target signal in the training dataset maintains high mutual information with a corresponding label and the information of spurious signals are reduced. We show training on the augmented datasets produces rationales on documents that better align with human annotation. Our experiments include six human produced datasets and two large-language model generated datasets.

数据增强反事实可解释性偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。