arXiv:2606.12316cs.CV2026-06被引 3

用可组合的视觉符号状态转换建模ARC规则,提升推理能力。

Slots, Transitions, Loops: Learning Composable World Models for ARC

论文配图:Slots, Transitions, Loops: Learning Composable World Models for ARC
图 1 · 摘自论文原文
  • 将视觉符号规则建模为对象状态间的可组合转换
  • 在ARC-1和ARC-2上超越基线,参数更少或相当
  • 适合研究视觉符号推理与世界模型的学者

ARC测试上下文规则归纳:给定若干输入-输出示例,模型需推断隐藏规则并应用于新查询。尽管已有方法通过语言、代码或符号程序表达ARC规则,但ARC本身是视觉符号的:规则表现为对象、颜色、形状和空间关系的网格变化。我们提出Loop-OWM,一种以对象为中心的世界模型架构,将这些规则学习为结构化状态上的可组合转换。它结合颜色原型槽、演示条件化的任务摘要,以及具有密集传播和槽条件修正的循环转换模型。在ARC-1和ARC-2上,Loop-OWM均优于非循环和循环基线,且参数量相当或更少。结果表明,ARC规则不仅可作为语言描述或搜索程序学习,还可作为视觉符号世界状态间的转换来建模。

原文摘要 · Abstract (English)

ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many approaches express ARC rules through language, code, or symbolic programs, ARC itself is visual-symbolic: rules appear as grid transitions over objects, colors, shapes, and spatial relations. We introduce Loop-OWM, an object-centric world-modeling architecture that learns these rules as composable transitions over structured states. It combines color-prototype slots, demonstration-conditioned task summaries, and a looped transition model with dense propagation and slot-conditioned correction. On both ARC-1 and ARC-2, Loop-OWM outperforms non-looped and looped baselines with comparable or fewer parameters. These results suggest that ARC rules can be learned not only as language descriptions or searched programs, but also as transitions over visual-symbolic world states.

视觉推理世界模型可组合性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。