对比GPT与o系列模型在多模态谜题上的推理演进,揭示其优势与瓶颈。
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
- 通过多模态谜题测试追踪GPT与o系列模型的推理能力演化。
- o3和o4-mini在视觉细粒度感知上显著优于GPT系列,具可扩展性。
- 仍难解决复杂组合逻辑或高精度视觉属性组合任务,指向AGI关键短板。
OpenAI推出的o-[n]系列(如o1、o3、o4-mini)标志着大语言模型向高级推理能力的重大转变。尽管o3在抽象与推理基准(ARC-AGI)上表现优异,但该基准仅限符号模式,而人类常基于视觉与语言的多模态信息进行推理。因此,亟需考察多模态任务中的先进推理能力。为此,我们评估GPT-[n]与o-[n]系列(含o1、o3、o4-mini)在PuzzleVQA与AlgoPuzzleVQA两个挑战性多模态谜题数据集上的表现,这些任务要求精细的视觉感知。结果显示,o-[n]系列(尤其是o3与o4-mini)显著优于GPT-[n]系列,并展现出强大的多模态推理可扩展性。然而,即便领先模型仍面临持续挑战,尤其在需要精确视觉感知、跨多个视觉属性的组合推理,以及解决复杂算法或高度组合性谜题的任务中,凸显未来AGI发展的关键方向。我们将持续跟踪系列新模型,并更新本研究结果。所有评估资源均公开于https://github.com/declare-lab/LLM-PuzzleTest。
原文摘要 · Abstract (English)
The releases of OpenAI's o-[n] series, such as o1, o3, and o4-mini, mark a significant paradigm shift in Large Language Models towards advanced reasoning capabilities. Notably, models like o3 have demonstrated strong performance on benchmarks like the Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI). However, this benchmark is limited to symbolic patterns, whereas humans often perceive and reason about multimodal scenarios involving both vision and language data. Thus, there is an urgent need to investigate advanced reasoning capabilities in multimodal tasks. To this end, we track the evolution of the GPT-[n] and o-[n] series models (including o1, o3, and o4-mini) on challenging multimodal puzzles from PuzzleVQA and AlgoPuzzleVQA, which demand fine-grained visual perception. Our results reveal that o-[n] series, particularly later iterations like o3 and o4-mini, significantly outperform the GPT-[n] series and show strong scalability in multimodal reasoning. Nonetheless, despite these substantial advancements and the superior capabilities demonstrated by the o-[n] series, our findings highlight that even these leading models face persistent challenges. Difficulties are particularly evident in tasks requiring precise visual perception, robust compositional reasoning across multiple visual attributes, and solving complex algorithmic or highly combinatorial puzzles, indicating critical areas for future AGI development. We plan to continuously track new models in the series and update our results in this paper accordingly. All resources used in this evaluation are openly available at https://github.com/declare-lab/LLM-PuzzleTest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。