对比人类与大模型在猜词游戏中的策略,发现模型存在词汇幻觉和重复问题。
Strategic Insights in Human and Large Language Model Tactics at Word Guessing Games
- 分析25%高频玩家的策略与动机,揭示长期参与行为特征
- 多模型在跨语言猜词中普遍存在长度错误与虚构词汇现象
- 适合关注LLM认知偏差与人类游戏行为差异的研究者
2022年初,一款简单的猜词游戏风靡全球,并被改编为多种语言版本。本文研究了超过两年时间内日常玩家策略的演变。通过对25%高频玩家的调查,揭示其持续参与的策略与动机。同时,评估了多个主流开源大语言模型在两种不同语言中理解与玩该游戏的能力。结果表明,部分模型难以准确维持猜测词的正确长度,频繁生成重复词,还存在虚构不存在的词汇及错误词形变化等幻觉现象。
原文摘要 · Abstract (English)
At the beginning of 2022, a simplistic word-guessing game took the world by storm and was further adapted to many languages beyond the original English version. In this paper, we examine the strategies of daily word-guessing game players that have evolved during a period of over two years. A survey gathered from 25% of frequent players reveals their strategies and motivations for continuing the daily journey. We also explore the capability of several popular open-access large language model systems and open-source models at comprehending and playing the game in two different languages. Results highlight the struggles of certain models to maintain correct guess length and generate repetitions, as well as hallucinations of non-existent words and inflections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。