用程序生成视觉推理题,测试大模型的逻辑能力
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- 用图形元素自动生成带真值答案的谜题
- 顶尖模型仅51.1%准确率,远低于人类表现
- 可验证奖励强化学习显著提升推理能力
我们提出Sphinx,一个面向视觉感知与推理的合成环境,聚焦核心认知能力。Sphinx通过图示、拼块、图表、图标和几何图形等元素程序化生成谜题,并配有可验证的真值解,支持精确评估与大规模数据集构建。该基准涵盖25种任务类型,包括对称性检测、几何变换、空间推理、图表解读和序列预测。对最新大视觉语言模型(LVLM)的评估显示,即使是最先进的GPT-5也仅达到51.1%准确率,远低于人类水平。最后,我们证明基于可验证奖励的强化学习(RLVR)能显著提升模型在这些任务上的表现,并在外部视觉推理基准上取得增益,展现出其在多模态推理中的潜力。
原文摘要 · Abstract (English)
We present Sphinx, a synthetic environment for visual perception and reasoning that targets core cognitive primitives. Sphinx procedurally generates puzzles using motifs, tiles, charts, icons, and geometric primitives, each paired with verifiable ground-truth solutions, enabling both precise evaluation and large-scale dataset construction. The benchmark covers 25 task types spanning symmetry detection, geometric transformations, spatial reasoning, chart interpretation, and sequence prediction. Evaluating recent large vision-language models (LVLMs) shows that even state-of-the-art GPT-5 attains only 51.1% accuracy, well below human performance. Finally, we demonstrate that reinforcement learning with verifiable rewards (RLVR) substantially improves model accuracy on these tasks and yields gains on external visual reasoning benchmarks, highlighting its promise for advancing multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。