arXiv:2507.11932cs.CV2025-07NeurIPS被引 5

构建新基准评估多模态大模型的内在想象能力

Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs

  • 设计四类程序化谜题,分三级难度测试模型内部构图能力
  • 顶尖模型在复杂任务上表现远低于人类,差距显著
  • 适合研究认知推理、视觉想象力的AI方向学者参考

心理可视化是人类认知的核心能力,涉及空间导航、物理轨迹预测和复杂视觉问题求解中的想象模拟。尽管多模态大模型(MLLMs)发展迅速,现有评测主要关注被动视觉感知,难以反映主动构建内部视觉表征的能力。为此,我们提出Hyperphantasia,一个合成基准,通过四个精心设计的谜题评估MLLM的心理可视化能力。每个谜题采用程序生成,包含三个难度层级,实现对模型性能在复杂度递增下的可控分析。对前沿模型的全面评估显示,人类与当前MLLMs之间存在显著性能差距。此外,我们探索了强化学习提升视觉模拟能力的可能性。结果表明,部分模型虽能识别视觉模式,但具备稳健心理可视化能力仍是未解挑战。

原文摘要 · Abstract (English)

Mental visualization, the ability to construct and manipulate visual representations internally, is a core component of human cognition and plays a vital role in tasks involving reasoning, prediction, and abstraction. Despite the rapid progress of Multimodal Large Language Models (MLLMs), current benchmarks primarily assess passive visual perception, offering limited insight into the more active capability of internally constructing visual patterns to support problem solving. Yet mental visualization is a critical cognitive skill in humans, supporting abilities such as spatial navigation, predicting physical trajectories, and solving complex visual problems through imaginative simulation. To bridge this gap, we introduce Hyperphantasia, a synthetic benchmark designed to evaluate the mental visualization abilities of MLLMs through four carefully constructed puzzles. Each puzzle is procedurally generated and presented at three difficulty levels, enabling controlled analysis of model performance across increasing complexity. Our comprehensive evaluation of state-of-the-art models reveals a substantial gap between the performance of humans and MLLMs. Additionally, we explore the potential of reinforcement learning to improve visual simulation capabilities. Our findings suggest that while some models exhibit partial competence in recognizing visual patterns, robust mental visualization remains an open challenge for current MLLMs.

多模态模型心理可视化认知评测程序生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。