提出新基准与几何方法,揭示视觉关系推理的通用机制。
Unraveling the geometry of visual relational reasoning
- 构建SimplifiedRPM基准,系统评估抽象关系推理能力。
- 发现SCL模型最接近人类表现,且存在信号与维度的权衡。
- 提出SNRloss损失函数,提升模型几何表示的合理性。
人类能轻松泛化抽象关系(如形状或颜色的恒定性),而神经网络则表现不佳,限制了其灵活推理能力。为探究此类泛化的内在机制,我们引入SimplifiedRPM——一个新型基准,用于系统评估抽象关系推理,弥补以往数据集的不足。同时开展人类实验,量化关系难度,实现模型与人类的直接对比。测试四种模型(ResNet-50、Vision Transformer、Wild Relation Network、Scattering Compositional Learner, SCL),发现SCL在泛化上表现最佳,且最贴近人类行为。通过几何分析,识别出可准确预测泛化性能的关键表示特性,并揭示信号与维度间的根本权衡:新关系被压缩至训练诱导的子空间中。层间分析揭示了关系结构的出现位置,指出瓶颈所在,并生成关于大脑抽象推理的可验证假设。基于这些洞见,我们提出SNRloss——一种显式平衡表示几何结构的新目标函数。研究建立了关系推理的几何基础,为实现更类人的视觉推理提供路径,并为将几何分析扩展至更广泛认知任务开辟前景。
原文摘要 · Abstract (English)
Humans readily generalize abstract relations, such as recognizing "constant" in shape or color, whereas neural networks struggle, limiting their flexible reasoning. To investigate mechanisms underlying such generalization, we introduce SimplifiedRPM, a novel benchmark for systematically evaluating abstract relational reasoning, addressing limitations in prior datasets. In parallel, we conduct human experiments to quantify relational difficulty, enabling direct model-human comparisons. Testing four models, ResNet-50, Vision Transformer, Wild Relation Network, and Scattering Compositional Learner (SCL), we find that SCL generalizes best and most closely aligns with human behavior. Using a geometric approach, we identify key representation properties that accurately predict generalization and uncover a fundamental trade-off between signal and dimensionality: novel relations compress into training-induced subspaces. Layer-wise analysis reveals where relational structure emerges, highlights bottlenecks, and generates concrete hypotheses about abstract reasoning in the brain. Motivated by these insights, we propose SNRloss, a novel objective explicitly balancing representation geometry. Our results establish a geometric foundation for relational reasoning, paving the way for more human-like visual reasoning in AI and opening promising avenues for extending geometric analysis to broader cognitive tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。