arXiv:2506.04668cs.CVcs.AI2025-06被引 2

用群分解理论提升无监督特征表示,更好模拟人类物体识别发展。

Feature-Based Lie Group Transformer for Real-World Applications

  • 将像素变换改为特征变换,结合目标分割实现更真实场景建模。
  • 在包含真实背景的复杂数据集上验证,可处理高阶条件独立性关系。
  • 适合研究人类认知发展与鲁棒视觉表征的科研人员。

表征学习的核心目标是从真实世界的感官输入中无监督地获取有意义的表征,该过程部分解释了人类认知发展机制。尽管已有多种神经网络模型能获得良好经验表征,但理想表征的定义尚未明确。我们此前提出一种通过代数结构约束学习双输入间变换的方法,发现传统假设的解耦独立特征轴无法解释条件独立性。为此,我们引入伽罗瓦代数中的群分解理论,但原方法依赖像素级对齐且仅适用于低分辨率、无背景图像,难以用于真实场景。本文提出新方法:用特征提取替代像素变换,将目标分割视为同质变换下的特征分组。我们在一个含真实物体与背景的实际数据集上验证了该方法的有效性,结果表明其能更合理建模现实世界中的表征学习过程,有望深化对人类物体识别发展的理解。

原文摘要 · Abstract (English)

The main goal of representation learning is to acquire meaningful representations from real-world sensory inputs without supervision. Representation learning explains some aspects of human development. Various neural network (NN) models have been proposed that acquire empirically good representations. However, the formulation of a good representation has not been established. We recently proposed a method for categorizing changes between a pair of sensory inputs. A unique feature of this approach is that transformations between two sensory inputs are learned to satisfy algebraic structural constraints. Conventional representation learning often assumes that disentangled independent feature axes is a good representation; however, we found that such a representation cannot account for conditional independence. To overcome this problem, we proposed a new method using group decomposition in Galois algebra theory. Although this method is promising for defining a more general representation, it assumes pixel-to-pixel translation without feature extraction, and can only process low-resolution images with no background, which prevents real-world application. In this study, we provide a simple method to apply our group decomposition theory to a more realistic scenario by combining feature extraction and object segmentation. We replace pixel translation with feature translation and formulate object segmentation as grouping features under the same transformation. We validated the proposed method on a practical dataset containing both real-world object and background. We believe that our model will lead to a better understanding of human development of object recognition in the real world.

表征学习群分解无监督学习物体识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。