用多生成器融合提升视觉语言合成数据多样性,显著优于单一模型生成。
PolyGen: Fully Synthetic Vision-Language Training via Multi-Generator Ensembles
- 采用多生成器集成构建合成数据,覆盖更广特征流形。
- 在多任务基准上比单源基线高19.0%,组合理解任务高9.1%。
- 适合关注合成数据质量与模型泛化性的研究者使用。
合成数据为视觉语言预训练提供可扩展的解决方案,但当前主流方法依赖单一生成模型,导致特定谱偏差并限制特征多样性。本文提出PolyGen框架,通过多生成器集成实现对特征流形的全面覆盖和结构化组合,有效消除模型特异性伪影。引入程序化困难负样本训练策略,强化细粒度语法理解。通过将相同数据预算从唯一描述转为多源变体,实现更鲁棒的特征空间。在综合多任务基准上超越领先单源基线SynthCLIP 19.0%,在SugarCrepe++组合性测试中提升9.1%。结果表明,结构多样性是比单纯增加数据量更高效的数据扩展规律。
原文摘要 · Abstract (English)
Synthetic data offers a scalable solution for vision-language pre-training, yet current state-of-the-art methods typically rely on scaling up a single generative backbone, which introduces generator-specific spectral biases and limits feature diversity. In this work, we introduce PolyGen, a framework that redefines synthetic data construction by prioritizing manifold coverage and compositional rigor over simple dataset size. PolyGen employs a Polylithic approach to train on the intersection of architecturally distinct generators, effectively marginalizing out model-specific artifacts. Additionally, we introduce a Programmatic Hard Negative curriculum that enforces fine-grained syntactic understanding. By structurally reallocating the same data budget from unique captions to multi-source variations, PolyGen achieves a more robust feature space, outperforming the leading single-source baseline (SynthCLIP) by +19.0% on aggregate multi-task benchmarks and on the SugarCrepe++ compositionality benchmark (+9.1%). These results demonstrate that structural diversity is a more data-efficient scaling law than simply increasing the volume of single-source samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。