用可调控的填字游戏评估大模型的多模态推理能力
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
- 通过可控生成机制设计文本与图像双格式填字谜题
- 推理型大模型在交叉字母约束下表现显著优于非推理模型
- 适合研究多模态推理、模型评估或对齐能力的研究者
现有大语言模型(LLM)和大视觉语言模型(LVLM)的推理评估框架大多仅聚焦于文本或视觉语言理解,缺乏文本与视觉约束间的动态交互。为此,我们提出CrossWordBench,一个基于填字谜题的基准测试,用于评估LLM与LVLM在语义线索与网格交叉约束双重约束下的推理能力。该基准采用可控生成框架,支持文本与图像两种格式,通过预填充率调节难度,并提供从直接求解到交互式求解等多种评估策略。对20余种模型的全面评估显示,推理型模型能有效利用交叉字母约束,表现显著更优;而LVLM在任务中表现受限,其解题能力与网格解析准确率高度相关。结果揭示了当前模型在多模态推理上的局限性,并为未来构建多模态约束任务提供了有效方法。
原文摘要 · Abstract (English)
Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly assess either text-based reasoning or vision-language understanding capabilities, with limited dynamic interplay between textual and visual constraints. To address this limitation, we introduce CrossWordBench, a benchmark designed to evaluate the reasoning capabilities of both LLMs and LVLMs through the medium of crossword puzzles -- a task requiring multimodal adherence to semantic constraints from text-based clues and intersectional constraints from visual grid structures. CrossWordBench leverages a controllable puzzle generation framework that produces puzzles in two formats (text and image), supports adjustable difficulty through prefill ratio control, and offers different evaluation strategies, ranging from direct puzzle solving to interactive modes. Our extensive evaluation of over 20 models reveals that reasoning LLMs substantially outperform non-reasoning models by effectively leveraging crossing-letter constraints. We further demonstrate that LVLMs struggle with the task, showing a strong correlation between their puzzle-solving performance and grid-parsing accuracy. Our findings highlight limitations of the reasoning capabilities of current LLMs and LVLMs, and provide an effective approach for creating multimodal constrained tasks for future evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。