发现大模型自循环训练会加剧偏见,提出用奖励拒采缓解。
Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop
- 构建自消耗绩效循环框架,模拟模型自我生成数据的迭代过程。
- 实验显示该循环使偏好偏差上升,差异性偏差下降。
- 提出基于奖励的拒采策略,提升自进化系统的可信度。
大型语言模型(LLMs)的快速发展推动了合成数据在训练新模型中的应用。然而,这形成了一个自我消耗的再训练循环:模型使用自身输出进行训练,可能导致性能下降并引发新兴偏见。在实际应用中,已部署的模型会通过用户反馈影响其生成的数据分布。例如,若某类用户长期被忽视,相关查询数据将减少。本文引入自消耗绩效循环(SCPL)概念,研究合成数据在受控绩效反馈下的偏见演化机制。由于真实生产系统中用户偏好数据难以获取,本研究在可控环境下隔离分析反馈驱动的偏见演变。聚焦两种循环模式——典型再训练与较少被研究的增量微调。在三个真实任务上的实验表明,绩效循环会增加偏好偏差,降低差异性偏差。为此,设计一种基于奖励的拒绝采样策略以减轻偏见,推动更可信的自改进系统发展。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has led to growing interest in using synthetic data to train future models. However, this creates a self-consuming retraining loop, where models are trained on their own outputs and may cause performance drops and induce emerging biases. In real-world applications, previously deployed LLMs may influence the data they generate, leading to a dynamic system driven by user feedback. For example, if a model continues to underserve users from a group, less query data will be collected from this particular demographic of users. In this study, we introduce the concept of \textbf{S}elf-\textbf{C}onsuming \textbf{P}erformative \textbf{L}oop (\textbf{SCPL}) and investigate the role of synthetic data in shaping bias during these dynamic iterative training processes under controlled performative feedback. This controlled setting is motivated by the inaccessibility of real-world user preference data from dynamic production systems, and enables us to isolate and analyze feedback-driven bias evolution in a principled manner. We focus on two types of loops, including the typical retraining setting and the incremental fine-tuning setting, which is largely underexplored. Through experiments on three real-world tasks, we find that the performative loop increases preference bias and decreases disparate bias. We design a reward-based rejection sampling strategy to mitigate the bias, moving towards more trustworthy self-improving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。