通过删减规则测试模型真推理能力,发现当前AI面对新规则大幅失准。
A Study of Rule Omission in Raven's Progressive Matrices
- 故意在训练中移除部分结构规则,检验模型泛化能力
- 变换模型在新规则下准确率骤降,暴露其依赖模式识别
- 揭示现有模型缺乏真正抽象推理能力,适合关注AI认知局限的研究者
类比推理是人类认知的核心,也是人工智能的长期挑战。雷文渐进矩阵(RPM)作为评估抽象推理能力的经典基准,要求模型推断隐藏的结构规则。尽管基于视觉和语言的模型在该任务上取得进展,但其表现是否反映真实推理能力仍存疑。本研究通过在训练中刻意省略部分结构规则,考察现代AI系统的泛化能力。对序列到序列的Transformer模型以及视觉架构CoPINet和双对比网络在不偏见雷文数据集(I-RAVEN)上的评估显示,虽然变压器在熟悉规则上表现优异,但在面对新规则或被省略规则时准确率显著下降。此外,标记级准确率与完整答案准确率之间的差距凸显了当前方法的根本缺陷。这些发现为深度学习模型的推理机制提供了新见解,强调需发展超越模式识别的鲁棒抽象推理架构。
原文摘要 · Abstract (English)
Analogical reasoning lies at the core of human cognition and remains a fundamental challenge for artificial intelligence. Raven's Progressive Matrices (RPM) serve as a widely used benchmark to assess abstract reasoning by requiring the inference of underlying structural rules. While many vision-based and language-based models have achieved success on RPM tasks, it remains unclear whether their performance reflects genuine reasoning ability or reliance on statistical shortcuts. This study investigates the generalization capacity of modern AI systems under conditions of incomplete training by deliberately omitting several structural rules during training. Both sequence-to-sequence transformer models and vision-based architectures such as CoPINet and the Dual-Contrast Network are evaluated on the Impartial-RAVEN (I-RAVEN) dataset. Experiments reveal that although transformers demonstrate strong performance on familiar rules, their accuracy declines sharply when faced with novel or omitted rules. Moreover, the gap between token-level accuracy and complete answer accuracy highlights fundamental limitations in current approaches. These findings provide new insights into the reasoning mechanisms underlying deep learning models and underscore the need for architectures that move beyond pattern recognition toward robust abstract reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。