arXiv:2604.21346cs.AIcs.CL2026-04被引 1

用符号输入揭示视觉推理中的表征瓶颈,发现抽象任务失败主因是表征而非推理。

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning

论文配图:Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
图 1 · 摘自论文原文
  • 将图像转为符号程序,用大语言模型进行结构化推理
  • 符号输入下准确率达90%以上,像素输入仍近随机水平
  • 适合研究视觉表征、多模态模型瓶颈的学者参考

视觉-语言模型在抽象视觉推理任务(如Bongard问题)上表现不佳,引发疑问:瓶颈在于推理还是表征?我们基于具有真实生成程序的合成基准Bongard-LOGO开展研究,对比直接处理原始图像的端到端视觉-语言模型与接收图像符号化输入的大语言模型(LLM)。通过符号输入作为诊断工具而非实际架构,提出分量-语法(C--G)范式,将Bongard-LOGO转化为基于LOGO风格动作程序或结构化描述的符号推理任务。结果表明,使用符号输入的LLM取得显著且一致的提升,在自由形式问题上准确率达到中90%级别,而强视觉基线在相同任务定义下仍接近随机水平。对输入格式、显式概念提示和最小视觉接地的消融实验显示,这些因素影响远小于从像素到符号结构的转变。研究确认表征是抽象视觉推理的关键瓶颈,并展示了符号输入如何作为受控的诊断上界。

原文摘要 · Abstract (English)

Vision--language models (VLMs) often fail on abstract visual reasoning benchmarks such as Bongard problems, raising the question of whether the main bottleneck lies in reasoning or representation. We study this on Bongard-LOGO, a synthetic benchmark of abstract concept learning with ground-truth generative programs, by comparing end-to-end VLMs on raw images with large language models (LLMs) given symbolic inputs derived from those images. Using symbolic inputs as a diagnostic probe rather than a practical multimodal architecture, our \emph{Componential--Grammatical (C--G)} paradigm reformulates Bongard-LOGO as a symbolic reasoning task based on LOGO-style action programs or structured descriptions. LLMs achieve large and consistent gains, reaching mid--90s accuracy on Free-form problems, while a strong visual baseline remains near chance under matched task definitions. Ablations on input format, explicit concept prompts, and minimal visual grounding show that these factors matter much less than the shift from pixels to symbolic structure. These results identify representation as a key bottleneck in abstract visual reasoning and show how symbolic input can serve as a controlled diagnostic upper bound.

视觉推理符号表示表征学习大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。