提出一种可扩展的图像分解方法,能清晰识别物体类别并高效建模。
Deep Sprite-based Image Models: An Analysis
- 基于精灵图的深度分解框架,提升图像结构解析能力
- 在CLEVR数据集上达到顶尖无监督分割效果,支持线性扩展
- 结果可解释性强,适合需要明确物体分类的场景
尽管基础模型推动了图像分割的进步,扩散模型也生成越来越逼真的图像,但从一组图像中识别重复模式这一看似简单的问题仍悬而未决。本文聚焦于基于精灵图(sprite-based)的图像分解模型,这类方法在聚类和图像分解方面展现出潜力,且因其高可解释性备受关注。然而,现有模型形式多样、需针对特定数据集定制,且难以扩展至多物体图像。本文深入分析其设计细节,识别核心组件,并在多个聚类基准上进行系统评估。基于此分析,提出一种深层精灵图图像分解方法,在标准CLEVR基准上性能媲美最先进无监督类别感知图像分割方法,能随物体数量线性扩展,显式识别物体类别,并以高度可解释的方式完整建模图像。
原文摘要 · Abstract (English)
While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple problem of identifying recurrent patterns in a collection of images remains very much open. In this paper, we focus on sprite-based image decomposition models, which have shown some promise for clustering and image decomposition and are appealing because of their high interpretability. These models come in different flavors, need to be tailored to specific datasets, and struggle to scale to images with many objects. We dive into the details of their design, identify their core components, and perform an extensive analysis on clustering benchmarks. We leverage this analysis to propose a deep sprite-based image decomposition method that performs on par with state-of-the-art unsupervised class-aware image segmentation methods on the standard CLEVR benchmark, scales linearly with the number of objects, identifies explicitly object categories, and fully models images in an easily interpretable way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。