多模型自消费训练中,人工筛选可能适得其反,导致长期对齐性下降。
When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

- 构建多模型自消费动态框架,分析系统收敛条件。
- 发现人工筛选在跨模型传播中会削弱甚至逆转对齐效果。
- 揭示协同训练中人工干预的潜在副作用,适合模型评估者参考。
基础模型正越来越多地使用前序模型生成的合成数据进行训练,而非仅依赖真实数据。这种自消费训练范式可能导致模型坍塌、发散或偏见放大。近期研究(Ferbach等,2024)表明,在循环中引入人工筛选可引导自消费模型趋向人类对齐行为,但这些分析局限于单一、孤立模型仅消费自身输出的场景。现实中,模型常交互并基于其他模型产生的输入-输出对进行训练。本文研究多模型环境下的自消费训练。我们首先形式化了相互作用的自消费模型框架,并刻画其动力学系统收敛至稳定点的条件。随后,分析人工筛选对单个模型自身对齐的影响(自影响),以及该影响如何在不同模型间传播(交叉影响)。与孤立模型中人工筛选始终提升对齐性不同,我们发现跨模型交互可能削弱甚至反转这一效应,最终损害长期对齐性。
原文摘要 · Abstract (English)
Foundation models are increasingly trained on synthetic data generated by prior model iterations rather than exclusively on real data. This self-consuming training paradigm can lead to model collapse, divergence, or bias amplification. Recent work (Ferbach et al., 2024) shows that incorporating human curation into the loop can steer a self-consuming model toward human-aligned behavior, but these analyses focus on a single, isolated model that solely consumes its own outputs. In practice, however, models often interact and train on input-output pairs produced by other models. This paper studies self-consuming training in the multi-model regime. We first formalize a framework for interacting self-consuming models and characterize when the resulting dynamical system converges to a stable point. We then examine how human curation of one model affects its own alignment (self-influence) and how such effects propagate to other models (cross-influence). Unlike isolated settings where human curation always enhances model alignment, we show that cross-model interactions can dampen or even invert this effect, ultimately degrading long-term alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。