arXiv:2604.20925cs.LG2026-04

用代数结构约束学习物体间关系,让模型像婴儿一样无监督理解世界

Unsupervised Learning of Inter-Object Relationships via Group Homomorphism

论文配图:Unsupervised Learning of Inter-Object Relationships via Group Homomorphism
图 1 · 摘自论文原文
  • 用群同态约束神经网络,分解像素变化为平移、形变等有意义成分
  • 在追逐/逃避场景中无标签分出多个物体,准确建模相对运动关系
  • 适合研究发育智能或可解释表示学习的学者参考

当前深度学习依赖海量数据中的统计相关性,与人类(尤其是婴幼儿)通过有限经验自主构建世界结构的能力形成鲜明对比。本文提出一种基于群运算层次关系的无监督表示学习方法,旨在模拟婴儿认知发展过程。该模型整合架构可同时完成物体分割与动态图像序列中的运动规律提取。通过引入代数中的同态作为神经网络的结构约束,模型将像素级变化结构性地分解为平移、形变等有意义的变换分量。基于发展科学发现的交互场景(追逐与躲避任务),实验表明模型可在无任何真值标注的情况下,成功将多个物体分割至独立槽位;同时,物体间的相对运动(如靠近或远离)被准确映射并组织到一维加性潜在空间中。结果表明,通过引入代数几何约束而非仅依赖统计相关性学习,可获得具有物理可解释性的“解耦表示”。本研究有助于理解婴幼儿如何内化环境规律为结构,并为构建具有发育智能的人工系统提供新视角。

原文摘要 · Abstract (English)

While current deep learning models achieve high performance by learning statistical correlations from vast datasets,which stands in stark contrast to human learning. They lack the flexibility of humans-particularly preverbal infants-to autonomously acquire the underlying structure of the world from limited experience and adapt to novel situations. In this study, we propose an unsupervised representation learning method based on a hierarchical relationship in group operations, rather than statistical independence, aiming to build a computational model of the cognitive development of infants. The proposed model features an integrated architecture that simultaneously performs object segmentation and the extraction of motion laws from dynamic image sequences. By introducing the Homomorphism from algebra as a structural constraint within a neural network, the model structurally separates pixel-level changes into meaningful, decomposed transformation components, such as translation and deformation. Using interaction scenes (chasing and evading tasks) based on developmental science findings, we experimentally demonstrate that the model can segment multiple objects into individual slots without any ground-truth labels. Furthermore, we confirmed that relative movements between objects, such as approaching or receding, are accurately mapped and structured into a one-dimensional additive latent space. These results suggest that by introducing algebraic geometric constraints rather than relying solely on statistical correlation learning, physically interpretable "disentangled representations" can be acquired. This study contributes to the understanding of the process by which infants internalize environmental laws as structures and provides a new perspective for constructing artificial systems with developmental intelligence.

无监督学习发育智能解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。