用猜词游戏测试大模型的语义理解与创意联想能力。
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
- 用猜词游戏设计多维度实验,考察模型生成线索和猜词表现。
- 模型在抽象词和模糊词上表现更差,反应其语义泛化短板。
- 适合评估大模型的隐喻推理与跨概念关联能力,对研发有参考价值。
本研究将猜词游戏(Codenames)作为评估工具,考察大语言模型(LLMs)在特定语言与认知技能上的表现。模型扮演游戏双方:一方生成涵盖多个目标词的提示词,另一方据此猜测目标词。通过控制词汇类型(抽象词与具体词、歧义词与单义词)或对手行为(快速或缓慢揭示词语),设计了多种实验场景。对比了近期商业及开源模型的表现,揭示了模型在策略选择、难点应对及能力局限方面的细节,为理解大模型在创造性语义关联任务中的实际表现提供了实证依据。
原文摘要 · Abstract (English)
This study utilizes the game Codenames as a benchmarking tool to evaluate large language models (LLMs) with respect to specific linguistic and cognitive skills. LLMs play each side of the game, where one side generates a clue word covering several target words and the other guesses those target words. We designed various experiments by controlling the choice of words (abstract vs. concrete words, ambiguous vs. monosemic) or the opponent (programmed to be faster or slower in revealing words). Recent commercial and open-weight models were compared side-by-side to find out factors affecting their performance. The evaluation reveals details about their strategies, challenging cases, and limitations of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。