arXiv:2602.01370cs.CVcs.AI2026-02被引 1

用多生成器融合提升视觉语言合成数据多样性,显著优于单一模型生成。

PolyGen: Fully Synthetic Vision-Language Training via Multi-Generator Ensembles

  • 采用多生成器集成构建合成数据,覆盖更广特征流形。
  • 在多任务基准上比单源基线高19.0%,组合理解任务高9.1%。
  • 适合关注合成数据质量与模型泛化性的研究者使用。

合成数据为视觉语言预训练提供可扩展的解决方案,但当前主流方法依赖单一生成模型,导致特定谱偏差并限制特征多样性。本文提出PolyGen框架,通过多生成器集成实现对特征流形的全面覆盖和结构化组合,有效消除模型特异性伪影。引入程序化困难负样本训练策略,强化细粒度语法理解。通过将相同数据预算从唯一描述转为多源变体,实现更鲁棒的特征空间。在综合多任务基准上超越领先单源基线SynthCLIP 19.0%,在SugarCrepe++组合性测试中提升9.1%。结果表明,结构多样性是比单纯增加数据量更高效的数据扩展规律。

原文摘要 · Abstract (English)

Synthetic data offers a scalable solution for vision-language pre-training, yet current state-of-the-art methods typically rely on scaling up a single generative backbone, which introduces generator-specific spectral biases and limits feature diversity. In this work, we introduce PolyGen, a framework that redefines synthetic data construction by prioritizing manifold coverage and compositional rigor over simple dataset size. PolyGen employs a Polylithic approach to train on the intersection of architecturally distinct generators, effectively marginalizing out model-specific artifacts. Additionally, we introduce a Programmatic Hard Negative curriculum that enforces fine-grained syntactic understanding. By structurally reallocating the same data budget from unique captions to multi-source variations, PolyGen achieves a more robust feature space, outperforming the leading single-source baseline (SynthCLIP) by +19.0% on aggregate multi-task benchmarks and on the SugarCrepe++ compositionality benchmark (+9.1%). These results demonstrate that structural diversity is a more data-efficient scaling law than simply increasing the volume of single-source samples.

合成数据多生成器视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。