用对抗协作生成高质量多模态数据,提升大模型实战能力
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
- 通过集体生成与对抗评判循环,自动合成多样且具挑战性的多模态数据
- 在多个基准上表现优于基线,验证了合成数据的有效性
- 适合研究多模态预训练、数据增强及生成式模型评估的从业者
本文旨在开发高效的自动化数据合成技术,以生成多模态训练数据,提升多模态大语言模型(MLLMs)解决复杂现实任务的能力。为此,我们提出一种新颖且通用的方法——集体对抗数据合成(CADS),通过集体智能确保生成数据的质量与多样性,并利用对抗学习生成具有挑战性的样本,有效推动模型改进。CADS包含两个循环阶段:集体对抗生成(CAD-Generate)与集体对抗判断(CAD-Judge),前者联合生成新颖多模态数据,后者协同评估生成数据质量。此外,引入对抗上下文优化机制,优化生成上下文以鼓励生成高价值、具挑战性的数据。基于CADS,我们构建了包含20,000条样本的MMSynthetic-20K数据集,并训练出R1-SyntheticVL模型,在多个基准测试中展现出优越性能。
原文摘要 · Abstract (English)
In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLLMs in solving complex real-world tasks. To this end, we propose Collective Adversarial Data Synthesis (CADS), a novel and general approach to synthesize high-quality, diverse and challenging multimodal data for MLLMs. The core idea of CADS is to leverage collective intelligence to ensure high-quality and diverse generation, while exploring adversarial learning to synthesize challenging samples for effectively driving model improvement. Specifically, CADS operates with two cyclic phases, i.e., Collective Adversarial Data Generation (CAD-Generate) and Collective Adversarial Data Judgment (CAD-Judge). CAD-Generate leverages collective knowledge to jointly generate new and diverse multimodal data, while CAD-Judge collaboratively assesses the quality of synthesized data. In addition, CADS introduces an Adversarial Context Optimization mechanism to optimize the generation context to encourage challenging and high-value data generation. With CADS, we construct MMSynthetic-20K and train our model R1-SyntheticVL, which demonstrates superior performance on various benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。