arXiv:2507.11153cs.CVcs.AI2025-07被引 1

评测大模型的色觉能力,发现其在颜色识别上的不足并提出改进方法。

Assessing Color Vision Test in Large Vision-language Models

  • 构建多类别、多难度的颜色视觉测试数据集
  • 分析大模型在颜色任务中的错误类型
  • 通过微调提升模型颜色识别能力,适合视觉-语言研究者

随着大型视觉-语言模型的广泛应用,其色觉能力至关重要。然而,这些模型的色觉表现尚未得到充分探索。为填补这一空白,我们定义了一项针对大型视觉-语言模型的颜色视觉测试任务,并构建了一个涵盖多种题型和难度级别的数据集。此外,我们分析了大型视觉-语言模型在颜色任务中的错误类型,并提出了针对性的微调策略,以提升其在颜色视觉测试中的表现。

原文摘要 · Abstract (English)

With the widespread adoption of large vision-language models, the capacity for color vision in these models is crucial. However, the color vision abilities of large visual-language models have not yet been thoroughly explored. To address this gap, we define a color vision testing task for large vision-language models and construct a dataset \footnote{Anonymous Github Showing some of the data https://anonymous.4open.science/r/color-vision-test-dataset-3BCD} that covers multiple categories of test questions and tasks of varying difficulty levels. Furthermore, we analyze the types of errors made by large vision-language models and propose fine-tuning strategies to enhance their performance in color vision tests.

视觉语言色觉评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。