arXiv:2410.04940cs.LGcs.CV2024-10被引 2

无需物体先验,视频预测让分布式模型学会可分离的物体表征

Next state prediction gives rise to entangled, yet compositional representations of objects

  • 用视频下一帧预测训练,让分布式模型自动学习物体可线性分离的表征
  • 分布式模型在下游任务中表现不输甚至优于带物体槽的模型
  • 部分重叠的神经编码能同时保持可分性与动态压缩优势,适合通用建模

组合表征被认为使人类能够泛化到组合爆炸的状态空间。具有可学习物体槽的模型通过将物体信息编码在独立潜变量中,在此类泛化上表现出潜力,但依赖强架构先验。而具有分布式表征的模型则使用重叠且可能纠缠的神经代码,其支持组合泛化的潜力仍待探索。本文研究分布式模型是否可通过无监督视频训练,发展出类似槽模型的线性可分物体表征。结果表明,令人惊讶的是,分布式模型在下游预测任务中常达到甚至超过槽模型的表现。此外,我们发现线性可分的物体表征可在无物体中心先验下出现,其中下一帧预测等辅助目标起关键作用。最后,我们观察到分布式模型的物体表征从未完全解耦,即使线性可分:多个物体可通过部分重叠的神经群体编码,仍能被线性分类器高效区分。我们推测,保留部分共享代码有助于分布式模型更好压缩物体动态,可能增强泛化能力。

原文摘要 · Abstract (English)

Compositional representations are thought to enable humans to generalize across combinatorially vast state spaces. Models with learnable object slots, which encode information about objects in separate latent codes, have shown promise for this type of generalization but rely on strong architectural priors. Models with distributed representations, on the other hand, use overlapping, potentially entangled neural codes, and their ability to support compositional generalization remains underexplored. In this paper we examine whether distributed models can develop linearly separable representations of objects, like slotted models, through unsupervised training on videos of object interactions. We show that, surprisingly, models with distributed representations often match or outperform models with object slots in downstream prediction tasks. Furthermore, we find that linearly separable object representations can emerge without object-centric priors, with auxiliary objectives like next-state prediction playing a key role. Finally, we observe that distributed models' object representations are never fully disentangled, even if they are linearly separable: Multiple objects can be encoded through partially overlapping neural populations while still being highly separable with a linear classifier. We hypothesize that maintaining partially shared codes enables distributed models to better compress object dynamics, potentially enhancing generalization.

表征学习视频预测分布式表征组合泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。