arXiv:2411.02671cs.LGcs.AI2024-11被引 6

用潜在概念变量提升大模型表格预测的公平性

Fair In-Context Learning via Latent Concept Variables

  • 通过潜变量学习敏感属性无关的特征,实现公平演示选择
  • 在多个表格数据集上,公平性指标显著优于传统启发式方法
  • 适合关注模型公平性的高风险应用开发者

大语言模型(LLMs)的上下文学习(ICL)能力使其广泛应用于不同数据类型的预测任务,包括表格数据。然而,在高风险场景中,其预训练数据中的社会偏见可能导致歧视性结果。本文研究了在表格数据上的上下文学习中存在的固有偏见问题,提出一种基于潜概念变量的最优演示选择方法,实现资源高效的任务适配。设计了降低预测结果与敏感变量相关性的数据增强策略,促进潜概念学习过程中的公平性。利用小型内部模型学习潜概念变量,并推广至大型外部模型。实验证明,该方法在多个表格数据集上相比多种启发式演示选择方法,显著提升了公平性表现。

原文摘要 · Abstract (English)

The emerging in-context learning (ICL) ability of large language models (LLMs) has prompted their use for predictive tasks in various domains with different data types, including tabular data, facilitated by serialization methods. However, with increasing applications in high-stakes domains, it has been shown that LLMs can inherit social bias and discrimination from their pre-training data. In this work, we investigate inherent bias in LLMs during in-context learning with tabular data. We focus on an optimal demonstration selection approach that utilizes latent concept variables for resource-efficient task adaptation. We design data augmentation strategies that reduce the correlation between predictive outcomes and sensitive variables, helping promote fairness during latent concept learning. We utilize the learned concept to select demonstrations and obtain fair predictions. The latent concept variables are learned using a smaller internal LLM and generalized to larger external LLMs. We empirically verify that the fair latent variable approach improves fairness results on tabular datasets compared to multiple heuristic demonstration selection methods.

大模型公平性表格数据潜变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。