arXiv:2511.01340cs.CVcs.CL2025-11被引 1

构建首个大规模图文谜题评测集,推动视觉语言模型理解创意隐喻能力

$\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles

  • 提出融合描述与代码化结构化推理的通用框架
  • 在1333道谜题上使模型准确率提升20%-30%
  • 适合研究多模态认知推理与创意解谜的学者使用

理解图文谜题(Rebus Puzzles)需要图像识别、常识推理、多步思维和基于图像的文字游戏等综合能力,对现有视觉语言模型仍是挑战。本文提出|∘→[BUS]|,一个包含1333个英文图文谜题的大规模多样基准,涵盖食品、习语、体育、金融、娱乐等18类主题,风格与难度各异。我们还设计了RebusDescProgICE框架,结合非结构化描述与代码化结构化推理,并采用基于推理的上下文示例选择策略,在闭源与开源模型上分别比链式思考方法提升2.1-4.1%和20-30%的性能。

原文摘要 · Abstract (English)

Understanding Rebus Puzzles (Rebus Puzzles use pictures, symbols, and letters to represent words or phrases creatively) requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, multi-step reasoning, image-based wordplay, etc., making this a challenging task for even current Vision-Language Models. In this paper, we present $\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$, a large and diverse benchmark of $1,333$ English Rebus Puzzles containing different artistic styles and levels of difficulty, spread across 18 categories such as food, idioms, sports, finance, entertainment, etc. We also propose $RebusDescProgICE$, a model-agnostic framework which uses a combination of an unstructured description and code-based, structured reasoning, along with better, reasoning-based in-context example selection, improving the performance of Vision-Language Models on $\left|\,\circlearrowright\,\boxed{\text{BUS}}\,\right|$ by $2.1-4.1\%$ and $20-30\%$ using closed-source and open-source models respectively compared to Chain-of-Thought Reasoning.

图文谜题多模态推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。