用可控合成数据微调视觉语言模型,显著提升真实场景表现。
Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
- 构建完全可控的数据生成与标注流程,消除偏差与分布不均。
- 仅130张合成数据即可实现均匀性能,较真实数据提升13%准确率。
- 适合追求高鲁棒性、低偏差的视觉语言模型研究者。
视觉语言模型(VLM)通过微调提升性能通常依赖于现实场景的随机数据收集与标注,但这一过程易引入偏差、错误和分布不平衡,导致过拟合与性能失衡。尽管已有研究尝试合成数据生成,但往往缺乏对数据分布和标注质量的控制。本文提出一种完全可控的数据生成与标注流程,获得无偏差、分布均衡且标注清晰的合成数据。以物体绝对位置的空间推理任务为例,我们在多个合成与真实世界基准上对先进VLM进行微调,并评估其在真实场景中的迁移能力。实验发现:1)在均衡数据上微调可实现视觉场景中一致的性能表现,仅需130个样本即可缓解常见偏差;2)在合成数据上微调使模型在真实数据(COCO)上的性能提升13%,优于在完整COCO训练集上微调的模型。
原文摘要 · Abstract (English)
Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of real-world scenes. Despite the improvements, this process is often prone to biases, errors, and distribution imbalance, resulting in overfitting and imbalanced performance. Although a few studies have explored synthetic data generation, they typically lack control over data distribution and annotation quality. In this work, we re-evaluate the potential of model fine-tuning by exploring a fully controlled data generation and annotation pipeline, obtaining bias-free data with balanced distribution and clean annotations. Using the spatial reasoning task of identifying the absolute position of an object as a use case, we fine-tune state-of-the-art VLMs and conduct exhaustive evaluations on both synthetic and real-world benchmarks, including transferability to real-world scenes. Our experiments reveal two key findings: 1) fine-tuning on balanced data yields uniform performance across the visual scene and mitigates common biases with as few as 130 samples; and 2) fine-tuning on synthetic stimuli improves performance by 13% on real-world data (COCO), outperforming models fine-tuned on the full COCO train set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。