arXiv:2510.09171cs.CV2025-10被引 2

无需真实图像,仅用领域名称即可生成多样化物体实例,提升细粒度识别性能。

Instance-Level Generation for Representation Learning

  • 基于领域名称合成跨域物体实例,解决标注数据稀缺问题。
  • 在7个细粒度识别基准上显著提升检索效果,超越现有方法。
  • 适合需要快速构建细粒度识别系统的研究者与开发者使用。

实例级识别(ILR)致力于识别个体对象而非宽泛类别,提供最细粒度的图像分类能力。然而,其细粒度特性使得大规模标注数据集的构建极为困难,限制了其在多领域中的实际应用。为此,我们提出一种新方法,可从多个领域在不同条件和背景下合成多样化的物体实例,形成大规模训练数据集。与以往自动数据合成工作不同,本方法首次在不依赖任何真实图像的情况下应对ILR特有挑战。在生成数据上微调基础视觉模型,显著提升了七个跨领域ILR基准上的检索性能。该方法为大规模数据收集与整理提供了高效、有效的替代方案,开创了一种新的ILR范式:仅需输入目标领域的名称即可实现端到端生成,解锁广泛的实际应用场景。代码与预训练模型已公开于https://github.com/yankungou/ILGen。

原文摘要 · Abstract (English)

Instance-level recognition (ILR) focuses on identifying individual objects rather than broad categories, offering the highest granularity in image classification. However, this fine-grained nature makes creating large-scale annotated datasets challenging, limiting ILR's real-world applicability across domains. To overcome this, we introduce a novel approach that synthetically generates diverse object instances from multiple domains under varied conditions and backgrounds, forming a large-scale training set. Unlike prior work on automatic data synthesis, our method is the first to address ILR-specific challenges without relying on any real images. Fine-tuning foundation vision models on the generated data significantly improves retrieval performance across seven ILR benchmarks spanning multiple domains. Our approach offers a new, efficient, and effective alternative to extensive data collection and curation, introducing a new ILR paradigm where the only input is the names of the target domains, unlocking a wide range of real-world applications. The code and pretrained models are publicly available at https://github.com/yankungou/ILGen.

细粒度识别数据生成视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。