让扩散模型学会按物体拆分图像视频,实现精准编辑。
Learning Object-Centric Representations Based on Slots in Real World Scenarios
- 用轻量级槽位机制注入对象条件,保持生成能力的同时提升可控性。
- 在图像和视频上实现最优的物体发现与无监督分割,支持删除/替换等操作。
- 适合需要精细控制生成内容的创作者、科研人员及交互式工具开发者。
人工智能的核心目标之一是将场景表示为离散物体的组合,以实现细粒度、可控制的图像和视频生成。然而,主流扩散模型以整体方式处理图像并依赖文本条件,与物体级编辑需求不匹配。本文提出一种框架,将强大的预训练扩散模型适配至物体中心合成,同时保留其生成能力。核心挑战在于平衡全局场景一致性与解耦的物体控制。方法通过轻量级槽位条件机制融入预训练模型,保留其视觉先验的同时提供物体特异性操控。针对图像,SlotAdapt引入注册标记用于背景/风格,以及槽位条件模块用于物体,降低文本条件偏差,在物体发现、分割、组合编辑和可控生成任务中达到当前最佳性能。进一步扩展至视频领域,采用不变槽位注意力(ISA)分离物体身份与姿态,并使用基于Transformer的时间聚合器,保持帧间物体表示与动态的一致性。该方法在无监督视频物体分割与重建上建立新基准,支持物体移除、替换和插入等高级编辑任务,无需显式监督。总体而言,本工作为图像与视频的物体中心生成建模提供了通用且可扩展的方案,推动了人机感知与机器学习的融合,拓展了创意、科学与实际应用中的交互式结构化生成工具设计空间。
原文摘要 · Abstract (English)
A central goal in AI is to represent scenes as compositions of discrete objects, enabling fine-grained, controllable image and video generation. Yet leading diffusion models treat images holistically and rely on text conditioning, creating a mismatch for object-level editing. This thesis introduces a framework that adapts powerful pretrained diffusion models for object-centric synthesis while retaining their generative capacity. We identify a core challenge: balancing global scene coherence with disentangled object control. Our method integrates lightweight, slot-based conditioning into pretrained models, preserving their visual priors while providing object-specific manipulation. For images, SlotAdapt augments diffusion models with a register token for background/style and slot-conditioned modules for objects, reducing text-conditioning bias and achieving state-of-the-art results in object discovery, segmentation, compositional editing, and controllable image generation. We further extend the framework to video. Using Invariant Slot Attention (ISA) to separate object identity from pose and a Transformer-based temporal aggregator, our approach maintains consistent object representations and dynamics across frames. This yields new benchmarks in unsupervised video object segmentation and reconstruction, and supports advanced editing tasks such as object removal, replacement, and insertion without explicit supervision. Overall, this work establishes a general and scalable approach to object-centric generative modeling for images and videos. By bridging human object-based perception and machine learning, it expands the design space for interactive, structured, and user-driven generative tools in creative, scientific, and practical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。