模仿人类认知发展,用少量数据训练多模态模型
Dreaming Out Loud: A Self-Synthesis Approach For Training Vision-Language Models With Developmentally Plausible Data
- 分四阶段迭代训练:从语言基础到视觉关联,再到自生成标注
- 在无标签图像上自动生成描述文本,混合真实与合成数据继续训练
- 适合对少样本多模态学习感兴趣的研究者或教育AI应用
当前大型语言模型虽能生成类人文本,但需海量数据训练。本文受人类认知发展启发,提出一种自合成训练方法,分四个阶段进行:第一阶段在小规模语料上从零训练语言能力;第二阶段将语言与视觉环境结合,使用带标签图像训练模型生成描述性标题;第三阶段为自合成阶段,模型为未标注图像生成标题,并混合合成与真实文本继续训练语言模块,扩展语言表达能力,类似人类对新体验的自我标注;第四阶段通过视觉问答与推理等任务发展高级认知技能。该方法验证了以发展性合理数据量训练多模态模型的可行性。
原文摘要 · Abstract (English)
While today's large language models exhibit impressive abilities in generating human-like text, they require massive amounts of data during training. We here take inspiration from human cognitive development to train models in limited data conditions. Specifically we present a self-synthesis approach that iterates through four phases: Phase 1 sets up fundamental language abilities, training the model from scratch on a small corpus. Language is then associated with the visual environment in phase 2, integrating the model with a vision encoder to generate descriptive captions from labeled images. In the "self-synthesis" phase 3, the model generates captions for unlabeled images, that it then uses to further train its language component with a mix of synthetic, and previous real-world text. This phase is meant to expand the model's linguistic repertoire, similar to humans self-annotating new experiences. Finally, phase 4 develops advanced cognitive skills, by training the model on specific tasks such as visual question answering and reasoning. Our approach offers a proof of concept for training a multimodal model using a developmentally plausible amount of data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。