仅用一张图就能发现重复元素,无需标注或先验知识。
Bottom-up Modeling of Repeated Elements via Single Image Analysis-by-Synthesis

- 通过重建目标学习图像空间原型,实现无监督的元素建模。
- 在FSC-147数据集上成功捕捉116张复杂图像中的结构一致性与类别内差异。
- 适合图像分析、自监督学习和少样本对象识别研究者参考。
我们解决从单张图像中发现重复元素的问题。与依赖大规模标注数据集、精选多图集合或物体分割掩码的现有方法不同,本文证明仅凭一张图像即可完全自底向上地学习有意义的对象模型,且不依赖除粗略尺度先验外的任何先验知识。我们的方法通过重建目标学习可调节的图像空间原型,使模型能够识别并合成同一图像中一致的对象实例。在来自FSC-147数据集的116张真实图像上的实验表明,该方法成功构建了连贯的元素模型,并捕捉到挑战性图像中的类别内变化。定性结果表明,相比经典分解、联合对齐及3D对象建模方法,本方法在重建质量和可解释性分解方面表现更优,同时保持简洁的二维形式。这些结果表明,仅凭单图学习即可涌现出有意义的对象发现能力。
原文摘要 · Abstract (English)
We address the problem of discovering repeated elements from a single image. In contrast to existing approaches that depend on large annotated datasets, curated multi-image collections, or object segmentation masks, we show that a single image can suffice to learn a meaningful object model in a completely bottom-up fashion, without any prior knowledge beyond a coarse scale prior. Our method learns a tunable image-space prototype of the repeated elements through a reconstruction objective, enabling the model to identify and synthesize consistent object instances within the same image. Experiments on 116 real images from the FSC-147 dataset demonstrate that our method successfully learns coherent element models and captures intra-category variation on challenging images. Qualitative results reveal superior reconstructions and interpretable decompositions compared to classical decomposition, joint alignment, and 3D object modeling methods, while maintaining a simple 2D formulation. These results suggest that meaningful object discovery can emerge from single image learning alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。