arXiv:2410.06841cs.CV2024-10中稿 · the Asian Conferen…被引 2

用大模型生成布局并合成图像,显著提升小样本目标检测性能。

Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis

  • 结合大模型推理与布局生成,自动扩展少样本标注空间。
  • 在COCO 5/10/30-shot设置下,mAP提升超140%/50%/35%。
  • 适合需要数据增强的少样本检测研究者使用。

扩散模型的进步使大量高质量数据生成成为可能,其中布局到图像合成(LIS)模型能根据空间布局(如边界框、掩码等)生成图像,具备生成真实感图像的能力,但布局遵循度有限。如何有效迁移此类模型以实现可扩展的小样本检测数据增强仍不明确。为此,我们提出一个协同框架,结合大语言模型(LLM)与LIS模型,超越现有生成式增强方法。利用LLM的推理能力,在仅有少量标注示例时生成新的边界框,推断标注空间的空间先验。同时引入新颖的布局感知CLIP分数用于样本排序,强化生成布局与图像间的耦合。在COCO少样本基准上取得显著提升:YOLOX-S基线在5/10/30-shot设置下mAP分别提升超过140%/50%/35%。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have enabled a wide range of works exploiting their ability to generate high-volume, high-quality data for use in various downstream tasks. One subclass of such models, dubbed Layout-to-Image Synthesis (LIS), learns to generate images conditioned on a spatial layout (bounding boxes, masks, poses, etc.) and has shown a promising ability to generate realistic images, albeit with limited layout-adherence. Moreover, the question of how to effectively transfer those models for scalable augmentation of few-shot detection data remains unanswered. Thus, we propose a collaborative framework employing a Large Language Model (LLM) and an LIS model for enhancing few-shot detection beyond state-of-the-art generative augmentation approaches. We leverage LLM's reasoning ability to extrapolate the spatial prior of the annotation space by generating new bounding boxes given only a few example annotations. Additionally, we introduce our novel layout-aware CLIP score for sample ranking, enabling tight coupling between generated layouts and images. Significant improvements on COCO few-shot benchmarks are observed. With our approach, a YOLOX-S baseline is boosted by more than 140%, 50%, 35% in mAP on the COCO 5-,10-, and 30-shot settings, respectively.

小样本检测布局生成大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。