研究多模态生成模型在自训练中如何衰退,提出缓解方案。
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
- 用多模态模型做递归生成-训练,模拟真实多智能体系统。
- 发现视觉语言对齐提升但图文任务方差增大,体现新衰减特征。
- 通过增加解码预算、模型多样性等方法有效缓解模型崩溃。
近期研究揭示了生成模型在持续使用自生成数据训练时性能逐步退化的风险。然而,现有研究仅限于单模态模型,难以反映真实场景中多种多模态智能体通过合成数据自主交互并持续演进的复杂情况。本文将合成数据训练与模型崩溃研究扩展至多模态视觉-语言生成系统,包括视觉-语言模型(VLMs)和文本到图像扩散模型,并考察多个模型间递归生成-训练循环的影响。结果表明,此前在单模态模型中观察到的模型崩溃在多模态背景下呈现不同特征,如视觉-语言对齐能力提升,同时在VLM图像描述任务上出现更大的方差。此外,我们发现增加解码预算、提升模型多样性以及使用冻结模型重标注等通用策略能有效缓解模型崩溃。研究成果为降低自进化多智能体系统中的模型崩溃风险,以及构建稳健的多模态合成数据集提供了初步洞察与实践指导。
原文摘要 · Abstract (English)
Recent research has highlighted the risk of generative model collapse, where performance progressively degrades when continually trained on self-generated data. However, existing exploration on model collapse is limited to single, unimodal models, limiting our understanding in more realistic scenarios, such as diverse multi-modal AI agents interacting autonomously through synthetic data and continually evolving. We expand the synthetic data training and model collapse study to multi-modal vision-language generative systems, such as vision-language models (VLMs) and text-to-image diffusion models, as well as recursive generate-train loops with multiple models. We find that model collapse, previously observed in single-modality generative models, exhibits distinct characteristics in the multi-modal context, such as improved vision-language alignment and increased variance in VLM image-captioning task. Additionally, we find that general approaches such as increased decoding budgets, greater model diversity, and relabeling with frozen models can effectively mitigate model collapse. Our findings provide initial insights and practical guidelines for reducing the risk of model collapse in self-improving multi-agent AI systems and curating robust multi-modal synthetic datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。