arXiv:2506.08191cs.CV2025-06

让模型学会拆解复杂场景中的多个物体,还能自动生成训练数据。

Generative Learning of Differentiable Object Models for Compositional Interpretation of Complex Scenes

  • 用可微渲染器分解物体的形状、颜色等属性,实现端到端训练。
  • 在多物体场景中重建效果优于现有方法,尤其擅长处理重叠物体。
  • 通过隐空间损失提升训练效率,适合做视觉理解与生成任务的研究者。

本研究基于解耦视觉先验(DVP)架构,该自编码器能将场景中物体分解为形状、大小、朝向和颜色等独立视觉属性,并以潜在参数控制可微渲染器进行图像重建,实现端到端梯度训练。本文扩展原DVP以支持场景中多个物体,利用其潜在表示的可解释性,通过解码器生成额外训练样本,并设计依赖图像空间与潜在空间联合损失的训练模式,显著缓解了图像空间重构损失中存在的大量平坦区域带来的训练难题。为评估性能,提出一个新基准,涵盖多个2D物体,比Multi-dSprites更具参数灵活性。与MONet和LIVE两个基线对比,结果表明该方法在重构质量及重叠物体分解能力上均更优。还分析了不同损失函数对梯度的影响,解释其训练效能机制,并讨论可微渲染在自编码器中的局限及其应对策略。

原文摘要 · Abstract (English)

This study builds on the architecture of the Disentangler of Visual Priors (DVP), a type of autoencoder that learns to interpret scenes by decomposing the perceived objects into independent visual aspects of shape, size, orientation, and color appearance. These aspects are expressed as latent parameters which control a differentiable renderer that performs image reconstruction, so that the model can be trained end-to-end with gradient using reconstruction loss. In this study, we extend the original DVP so that it can handle multiple objects in a scene. We also exploit the interpretability of its latent by using the decoder to sample additional training examples and devising alternative training modes that rely on loss functions defined not only in the image space, but also in the latent space. This significantly facilitates training, which is otherwise challenging due to the presence of extensive plateaus in the image-space reconstruction loss. To examine the performance of this approach, we propose a new benchmark featuring multiple 2D objects, which subsumes the previously proposed Multi-dSprites dataset while being more parameterizable. We compare the DVP extended in these ways with two baselines (MONet and LIVE) and demonstrate its superiority in terms of reconstruction quality and capacity to decompose overlapping objects. We also analyze the gradients induced by the considered loss functions, explain how they impact the efficacy of training, and discuss the limitations of differentiable rendering in autoencoders and the ways in which they can be addressed.

可微渲染物体分解生成模型自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。