arXiv:2603.02063cs.CV2026-03

用对抗生成网络实现物体中心表征学习,能处理复杂真实图像。

ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks

  • 基于循环一致GAN构建物体中心表征框架
  • 在真实数据集上表现优于现有方法,支持多物体低对比度场景
  • 可扩展至大图和多物体,适合需要物体操作的下游任务

尽管数据生成相对简单,但从数据中提取信息更具挑战性。物体中心表征学习可无监督地将图像分解为独立物体,并在低维潜在空间中表示每个物体,便于后续处理。当前主流方法依赖自编码器架构(AEs)。本文提出新型方法 ORGAN,基于循环一致生成对抗网络实现物体中心表征学习。实验表明,ORGAN 在合成数据集上性能与最先进方法相当;同时,它是唯一能在包含多个物体且视觉对比度低的真实数据集上有效工作的方法。此外,ORGAN 构建的潜在空间表达能力强,支持物体操控。最后,其在物体数量和图像尺寸方面均具有良好可扩展性,相较当前最优方法具有独特优势。

原文摘要 · Abstract (English)

Although data generation is often straightforward, extracting information from data is more difficult. Object-centric representation learning can extract information from images in an unsupervised manner. It does so by segmenting an image into its subcomponents: the objects. Each object is then represented in a low-dimensional latent space that can be used for downstream processing. Object-centric representation learning is dominated by autoencoder architectures (AEs). Here, we present ORGAN, a novel approach for object-centric representation learning, which is based on cycle-consistent Generative Adversarial Networks instead. We show that it performs similarly to other state-of-the-art approaches on synthetic datasets, while at the same time being the only approach tested here capable of handling more challenging real-world datasets with many objects and low visual contrast. Complementing these results, ORGAN creates expressive latent space representations that allow for object manipulation. Finally, we show that ORGAN scales well both with respect to the number of objects and the size of the images, giving it a unique edge over current state-of-the-art approaches.

物体中心生成模型表征学习对抗网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。