arXiv:2505.09768cs.LG2025-05ICML被引 9

研究生成模型在虚假数据循环训练中的稳定性,发现恶意数据篡改可导致模型偏离用户真实偏好。

Self-Consuming Generative Models with Adversarially Curated Data

  • 基于对抗性数据筛选的自循环训练机制,分析模型演化规律。
  • 实验表明,少量恶意数据可显著误导模型分布,偏差达30%以上。
  • 适用于模型安全评估与对抗攻防场景,尤其关注平台间竞争风险。

生成模型的发展使得真实数据与合成数据难以区分。利用合成数据持续训练下一代模型会形成“自消耗循环”,可能导致模型崩溃或训练不稳定。尽管已有研究指出,若数据按用户偏好筛选,模型将收敛至优化该偏好的分布,但现实中数据筛选常含噪声或遭恶意操纵。例如,竞争对手可能招募恶意用户对数据进行对抗性筛选,以破坏对方模型。本文研究在噪声及对抗性数据筛选下的自消耗重训练过程,理论上分析其对生成模型的影响,并识别重训练过程鲁棒性的条件。基于此,设计了针对有限预算平台的攻击算法,通过恶意用户使对手模型偏离真实用户偏好。在合成与真实数据集上的实验验证了所提算法的有效性。

原文摘要 · Abstract (English)

Recent advances in generative models have made it increasingly difficult to distinguish real data from model-generated synthetic data. Using synthetic data for successive training of future model generations creates "self-consuming loops", which may lead to model collapse or training instability. Furthermore, synthetic data is often subject to human feedback and curated by users based on their preferences. Ferbach et al. (2024) recently showed that when data is curated according to user preferences, the self-consuming retraining loop drives the model to converge toward a distribution that optimizes those preferences. However, in practice, data curation is often noisy or adversarially manipulated. For example, competing platforms may recruit malicious users to adversarially curate data and disrupt rival models. In this paper, we study how generative models evolve under self-consuming retraining loops with noisy and adversarially curated data. We theoretically analyze the impact of such noisy data curation on generative models and identify conditions for the robustness of the retraining process. Building on this analysis, we design attack algorithms for competitive adversarial scenarios, where a platform with a limited budget employs malicious users to misalign a rival's model from actual user preferences. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed algorithms.

生成模型对抗攻击数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。