arXiv:2411.02116cs.CLcs.CV2024-11被引 3

对比人类与大模型在颜色-词汇联想上的表现,发现大模型仍有明显局限。

Advancements and limitations of LLMs in replicating human color-word associations

  • 用10,000人数据对比GPT-3到GPT-4o在17色80词上的表现
  • GPT-4o准确率最高但中位数仅约50%,远低于人类
  • 情绪类词汇关联能力差,反映语义记忆差异

颜色-词汇关联在人类认知与设计应用中具有基础作用。尽管大型语言模型(LLMs)在各类基准测试中表现出自然对话能力,其复制人类颜色-词汇关联的能力仍缺乏研究。本研究基于超过10,000名日本参与者的数据,对比了从GPT-3到GPT-4o多个世代的LLMs在17种颜色和80个词汇(来自八个类别)上的表现。结果表明,模型性能随代际提升,GPT-4o在预测每种颜色与类别最佳词汇时准确率最高。然而,即使使用视觉输入,其最高中位性能仍约为50%(随机水平为10%)。模型在节奏、景观等类别表现较好,但在情绪类词汇上表现不佳。通过颜色-词汇关联数据估算的颜色辨别能力与人类模式高度一致,说明大模型虽具备基本颜色辨别能力,但在词汇选择上仍系统性偏离人类。研究揭示了大模型在语义记忆结构上与人类的潜在差异。

原文摘要 · Abstract (English)

Color-word associations play a fundamental role in human cognition and design applications. Large Language Models (LLMs) have become widely available and have demonstrated intelligent behaviors in various benchmarks with natural conversation skills. However, their ability to replicate human color-word associations remains understudied. We compared multiple generations of LLMs (from GPT-3 to GPT-4o) against human color-word associations using data collected from over 10,000 Japanese participants, involving 17 colors and 80 words (10 word from eight categories) in Japanese. Our findings reveal a clear progression in LLM performance across generations, with GPT-4o achieving the highest accuracy in predicting the best voted word for each color and category. However, the highest median performance was approximately 50% even for GPT-4o with visual inputs (chance level of 10%). Moreover, we found performance variations across word categories and colors: while LLMs tended to excel in categories such as Rhythm and Landscape, they struggled with categories such as Emotions. Interestingly, color discrimination ability estimated from our color-word association data showed high correlation with human color discrimination patterns, consistent with previous studies. Thus, despite reasonable alignment in basic color discrimination, humans and LLMs still diverge systematically in the words they assign to those colors. Our study highlights both the advancements in LLM capabilities and their persistent limitations, raising the possibility of systematic differences in semantic memory structures between humans and LLMs in representing color-word associations.

大模型颜色感知语义记忆认知研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。