arXiv:2412.03927cs.CVcs.LG2024-12AAAI被引 2

构建高质量图像数据集,提升视觉语言模型对颜色与环境的感知能力。

MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models

  • 基于22万张真实图像构建多标签数据集,标注前景色、背景色及环境描述。
  • 用该数据集微调的小模型可超越闭源GPT-4o在颜色识别上的表现。
  • 适用于评估模型跨域泛化能力,推动视觉语言模型向更精细理解演进。

在视觉语言模型中,准确感知和理解颜色与物理环境对于实现情境化理解与可靠应用至关重要。然而,现有多模态模型仍缺乏专门评测其细微色彩差异与空间上下文辨别能力的数据集。为此,我们构建了高质量、人工标注的MegaCOIN数据集,基于真实图像并包含多种上下文属性。MegaCOIN由两部分组成:用于监督微调(SFT)的MegaCOIN-Instruct,以及可独立使用的测试集MegaCOIN-Bench。该数据集涵盖22万张真实图像,共提供66万条人工标注,包括前景色、背景色及物体所处物理环境的描述。此外,MegaCOIN可用于评估领域泛化(DG)算法。我们在线性探测设置下测试了多种DG方法,揭示新见解。实验表明,包括GPT-4o在内的现有视觉语言模型在颜色识别上表现不佳;而使用MegaCOIN微调后的开源小模型(如LLaVA、Bunny)在部分视觉评估任务中已优于闭源的GPT-4o。我们希望MegaCOIN能为视觉语言模型改进方向提供参考,并成为领域泛化算法的更复杂评测平台。

原文摘要 · Abstract (English)

In vision-language models (VLMs), the ability to perceive and interpret color and physical environment is crucial for achieving contextually accurate understanding and interaction. However, despite advances in multimodal modeling, there remains a significant lack of specialized datasets that rigorously evaluate a model's capacity to discern subtle color variations and spatial context -- critical elements for situational comprehension and reliable deployment across real-world applications. Toward that goal, we curate MegaCOIN, a high-quality, human-labeled dataset based on \emph{real} images with various contextual attributes. MegaCOIN consists of two parts: MegaCOIN-Instruct, which serves as a supervised fine-tuning (SFT) dataset for VLMs; and MegaCOIN-Bench, an annotated test set that can be used as a stand-alone QA dataset. MegaCOIN~provides three annotated features for 220,000 real images: foreground color, background color, and description of an object's physical environment, constituting 660k human annotations. In addition, MegaCOIN can be applied to benchmark domain generalization (DG) algorithms. We explore benchmarking DG methods in the linear probing setup for VLM and show some new insights. Last but not least, we show that VLMs, including GPT-4o, have subpar color recognition capabilities, and fine-tuning with MegaCOIN can result in improved performance on visual evaluation tasks. In certain cases, MegaCOIN fine-tuned small-scale opensource models such as LLaVA and Bunny can outperform closed-source GPT-4o. We hope the utilities of MegaCOIN can shed light on the directions VLMs can improve and provide a more complex platform for domain generalization algorithms.

视觉语言模型颜色感知数据集构建域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。