arXiv:2511.23450cs.CV2025-11

用少量物体数据生成逼真图像,提升新类别检测性能。

Object-Centric Data Synthesis for Category-level Object Detection

  • 基于多视角图或3D模型合成带复杂背景的图像。
  • 在数据受限条件下,检测性能显著提升。
  • 适合缺乏标注数据的新类别检测任务。

深度学习在特定物体类别的检测上已取得可靠成果。然而,将模型扩展到新类别需要大量标注数据,获取成本高,尤其对长尾类别而言更为困难。本文提出物体中心数据设置,当仅有少量物体中心数据(如多视角图像或3D模型)时,系统评估四种不同数据合成方法在该设定下微调目标检测模型以识别新类别物体的表现。这些方法包括简单图像处理、3D渲染和图像扩散模型,利用物体中心数据生成具有不同上下文一致性和复杂度的真实感、杂乱图像。我们评估了这些方法如何帮助模型在真实世界数据中实现类别级泛化,并证明在此数据受限设定下性能有显著提升。

原文摘要 · Abstract (English)

Deep learning approaches to object detection have achieved reliable detection of specific object classes in images. However, extending a model's detection capability to new object classes requires large amounts of annotated training data, which is costly and time-consuming to acquire, especially for long-tailed classes with insufficient representation in existing datasets. Here, we introduce the object-centric data setting, when limited data is available in the form of object-centric data (multi-view images or 3D models), and systematically evaluate the performance of four different data synthesis methods to finetune object detection models on novel object categories in this setting. The approaches are based on simple image processing techniques, 3D rendering, and image diffusion models, and use object-centric data to synthesize realistic, cluttered images with varying contextual coherence and complexity. We assess how these methods enable models to achieve category-level generalization in real-world data, and demonstrate significant performance boosts within this data-constrained experimental setting.

目标检测数据合成小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。