用Minecraft猜谜游戏评测视觉语言模型的创意理解能力。
GuessBench: Sensemaking Multimodal Creativity in the Wild
- 基于真实玩家游戏数据构建多模态猜谜任务
- GPT-4o在34%案例中出错,开源模型平均准确率仅13.87%
- 适合研究创意理解、跨文化推理和模型偏见的学者
我们提出GuessBench,一个评估视觉语言模型(VLMs)在真实世界中建模普遍性、噪声性和多元性人类创造力的新基准。该基准源自在线多人Minecraft小游戏「Guess the Build」,一名玩家根据概念(如毛毛虫)建造结构,其他玩家通过自然语言提示猜测其内容,为VLM作为猜测者提供了纯净的创意理解测试场景。我们从实际游戏过程中收集了1500张图像,设计了2000个问题,涵盖静态与动态图像设置、不同完整度的语言提示等。对六种开放/API VLM及五种推理增强方法的广泛实验表明,GuessBench提出了极具挑战性的创造力建模任务:即使最先进的GPT-4o在34%的实例中出错,且开放模型与API模型平均准确率差距高达13.87%对比53.93%。将GuessBench的推理轨迹用于微调,可使视觉感知任务平均提升15.36%。进一步分析发现,模型在创意理解中的表现与概念在训练数据中的频率相关,而在低资源语言和代表性不足的文化语境下准确率急剧下降。
原文摘要 · Abstract (English)
We propose GuessBench, a novel benchmark that evaluates Vision Language Models (VLMs) on modeling the pervasive, noisy, and pluralistic human creativity. GuessBench sources data from "Guess the Build", an online multiplayer Minecraft minigame where one player constructs a Minecraft build given a concept (e.g. caterpillar) and others try to guess it with natural language hints, presenting a pristine testbed for sensemaking creativity in the wild with VLMs acting as guessers. We curate 1500 images from the actual gameplay and design 2000 problems spanning static and dynamic image settings, natural language hints of varying completeness, and more. Extensive experiments with six open/API VLMs and five reasoning enhancement approaches demonstrate that GuessBench presents a uniquely challenging task in creativity modeling: even the start-of-the-art GPT-4o is incorrect on 34% of instances, while we observe a huge performance gap (13.87% vs. 53.93% on average) between open and API models. When used as a resource to improve VLMs, fine-tuning on the reasoning traces for GuessBench problems improves visual perception tasks by 15.36% on average. Further analysis reveals that VLM performance in creativity sensemaking correlates with the frequency of the concept in training data, while the accuracy drops sharply for concepts in underrepresented cultural contexts and low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。