将视觉语言模型缩小10倍仍能生成连贯文本,借鉴幼儿学习方式。
NanoVLMs: How small can we go and still make coherent Vision Language Models?
- 用儿童语言风格数据集训练极小模型,提升理解与表达能力。
- 模型体积仅是当前最小VLM的十分之一,仍能生成流畅一致文本。
- 通过人工评分多维度评估输出质量,适配资源受限场景使用。
视觉语言模型(VLMs)如GPT-4V和Llama 3.2 vision在多模态任务中表现出色,但受限于专有技术、高算力需求及可访问性差。现有小型模型如GIT和BLIP在长文本生成上表现不佳,难以维持连贯性。本文受3-4岁儿童依赖视觉理解与表达的启发,构建两个新数据集:ShortDesc(简短描述)与LongDesc(详细描述),均采用儿童常用词汇与语法,由缩小版GPT-4o生成。利用这些数据,我们成功训练出比当前最先进小规模VLM小10倍的模型,同时保持架构简洁。为评估输出,使用GPT-4o以学生作文标准对创造力、意义性与一致性进行评分(满分10分),克服传统基准在非结构化输出上的局限,实现多维能力评估。研究成果推动轻量、易部署的多模态模型发展,适用于资源受限环境。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs), such as GPT-4V and Llama 3.2 vision, have garnered significant research attention for their ability to leverage Large Language Models (LLMs) in multimodal tasks. However, their potential is constrained by inherent challenges, including proprietary restrictions, substantial computational demands, and limited accessibility. Smaller models, such as GIT and BLIP, exhibit marked limitations, often failing to generate coherent and consistent text beyond a few tokens, even with extensive training. This underscores a pivotal inquiry: how small can a VLM be and still produce fluent and consistent text? Drawing inspiration from the exceptional learning process of 3-4 year old children, who rely heavily on visual cues for understanding and communication, we introduce two novel datasets: ShortDesc (featuring concise image descriptions) and LongDesc (containing more detailed image descriptions). These datasets consist of image-text pairs where the text is restricted to the simple vocabulary and syntax typically used by young children, generated with a scaled-down model, GPT-4o. Using these datasets, we demonstrate that it is possible to train VLMs that are significantly smaller, up to 10 times smaller than state of the art(SOTA) small VLMs while maintaining architectural simplicity. To evaluate the outputs, we leverage GPT-4o to grade the text, as if stories written by students, on creativity, meaningfulness, and consistency, assigning scores out of 10. This method addresses limitations of standard benchmarks by accommodating unstructured outputs and providing a multidimensional evaluation of the model capabilities. Our findings contribute to the development of lightweight, accessible multimodal models for resource constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。