arXiv:2509.02305cs.CV2025-09被引 1

用桌游测试CLIP的色彩感知能力,发现其与人类高度一致但存在文化偏差。

Hues and Cues: Human vs. CLIP

  • 用桌游Hues & Cues测试CLIP的色彩识别与命名能力。
  • CLIP在多数情况下与人类表现接近,但在抽象层次上存在不一致。
  • 桌游评测可暴露传统基准难以发现的模型缺陷,适合评估人机对齐。

游戏是人类固有的行为,许多游戏旨在挑战人类的不同认知特征。然而,这些任务在评估人工模型的人类相似性时常被忽视。本文提出一种新方法,通过玩桌游Hues & Cues来评估人工模型的性能。我们测试了CLIP在色彩感知和命名方面的能力,并评估其与人类观察者的对齐程度。实验表明,CLIP整体上与人类观察者高度一致,但该方法揭示了模型在处理不同抽象层级时存在的文化偏见和不一致性,这些不足在其他测试策略中难以察觉。研究指出,通过如桌游等多样化任务评估模型,能够以传统基准难以实现的方式凸显模型缺陷。

原文摘要 · Abstract (English)

Playing games is inherently human, and a lot of games are created to challenge different human characteristics. However, these tasks are often left out when evaluating the human-like nature of artificial models. The objective of this work is proposing a new approach to evaluate artificial models via board games. To this effect, we test the color perception and color naming capabilities of CLIP by playing the board game Hues & Cues and assess its alignment with humans. Our experiments show that CLIP is generally well aligned with human observers, but our approach brings to light certain cultural biases and inconsistencies when dealing with different abstraction levels that are hard to identify with other testing strategies. Our findings indicate that assessing models with different tasks like board games can make certain deficiencies in the models stand out in ways that are difficult to test with the commonly used benchmarks.

模型评估视觉理解人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。