arXiv:2602.03038cs.CVcs.AI2026-02

用程序化规则+大模型解决视觉推理难题

Bongards at the Boundary of Perception and Reasoning: Programs or Language?

  • 将视觉推理规则转化为可执行程序,由大模型生成并优化
  • 在已知规则下准确分类图像,零样本解题成功率超基准方法
  • 适合研究视觉推理与神经符号系统融合的学者

视觉语言模型在日常视觉任务中表现优异,如自然图像描述或常识问答。但人类具备在全新情境中运用视觉推理的能力,这通过经典的Bongard问题得到严格检验。本文提出一种神经符号方法:给定一个假设的Bongard问题解题规则,利用大语言模型生成该规则的参数化程序表示,并通过贝叶斯优化进行参数拟合。我们在已知真实规则条件下评估该方法的图像分类能力,并在无先验知识情况下从头求解问题。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual reasoning abilities in radically new situations, a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parameter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.

视觉推理神经符号大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。