让AI像写程序一样理解视觉集合推理,更准更透明。
Visual Set Program Synthesizer
- 先生成可执行的符号程序,再在视觉场景中运行
- 在新基准Set-VQA上准确率显著超越现有模型
- 适合需要可解释推理的复杂视觉任务
用户用手机对准超市货架询问“哪种汽水糖分最少?”这类问题,要求不仅识别物体,还需进行过滤、比较和聚合等集合推理。当前端到端多模态大模型因缺乏显式组合逻辑机制,常无法胜任此类任务。本文提出将视觉推理视为视觉程序合成,即模型先生成一个符号程序,由独立引擎在视觉场景中执行。同时引入专为评估集合推理设计的新基准Set-VQA。实验表明,该方法在复杂推理任务上显著优于现有基线,行为更系统、透明,答案准确率大幅提升。结果证明,程序驱动推理为黑箱视觉语言推断提供了一种更合理的替代方案。
原文摘要 · Abstract (English)
A user pointing their phone at a supermarket shelf and asking "Which soda has the least sugar?" poses a difficult challenge for current visual Al assistants. Such queries require not only object recognition, but explicit set-based reasoning such as filtering, comparison, and aggregation. Standard endto-end MLLMs often fail at these tasks because they lack an explicit mechanism for compositional logic. We propose treating visual reasoning as Visual Program Synthesis, where the model first generates a symbolic program that is executed by a separate engine grounded in visual scenes. We also introduce Set-VQA, a new benchmark designed specifically for evaluating set-based visual reasoning. Experiments show that our approach significantly outperforms state-of-the-art baselines on complex reasoning tasks, producing more systematic and transparent behavior while substantially improving answer accuracy. These results demonstrate that program-driven reasoning provides a principled alternative to black-box visual-language inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。