让两个AI玩图像识别游戏,自动生成高质量训练数据。
Vision-Language Model Dialog Games for Self-Improvement
- 两个AI互相出题识别图像,通过成功互动筛选优质数据。
- 用自动生成的数据微调后,模型在多个任务上表现更好。
- 适合缺乏高质量多模态数据的场景,可反复迭代提升。
高质量、多样化的训练数据日益成为推动视觉语言模型(VLMs)发展的瓶颈。本文提出一种名为VLM Dialog Games的新颖且可扩展的自改进框架。该方法通过两个智能体围绕图像识别目标进行目标导向的交互对弈,利用成功的游戏交互自动构建高质量的图像与文本交错数据集。实验表明,在此合成数据上微调模型能显著提升下游任务性能,并实现跨数据集泛化。更重要的是,随着模型能力提升,游戏表现也随之改善,使该过程可循环迭代。本工作为自进化视觉语言模型开辟了新路径,尤其适用于高质量多模态数据稀缺的实际场景。
原文摘要 · Abstract (English)
The increasing demand for high-quality, diverse training data poses a significant bottleneck in advancing vision-language models (VLMs). This paper presents VLM Dialog Games, a novel and scalable self-improvement framework for VLMs. Our approach leverages self-play between two agents engaged in a goal-oriented play centered around image identification. By filtering for successful game interactions, we automatically curate a high-quality dataset of interleaved images and text. We demonstrate that fine-tuning on this synthetic data leads to performance gains on downstream tasks and generalises across datasets. Moreover, as the improvements in the model lead to better game play, this procedure can be applied iteratively. This work paves the way for self-improving VLMs, with potential applications in various real-world scenarios especially when the high-quality multimodal data is scarce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。