arXiv:2604.02966cs.CV2026-04中稿 · ed

用视觉原型提升无人机检测生成图像质量,减少边界伪影。

Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection

  • 以视觉原型引导扩散模型生成更真实的物体
  • 聚焦前景区域并修正标注错误,提升小物体生成精度
  • 适用于数据少、场景变化大的无人机目标检测任务

基于无人机的目标检测在动态变化场景中面临标注数据有限的挑战。布局到图像生成方法通过扩散模型合成带标签图像,能提升检测准确率,但常在小物体边界产生伪影,严重影响性能。为此,我们提出UAVGen框架,包含视觉原型条件扩散模型(VPC-DM),为每类构建代表性实例并融入潜在嵌入,实现高保真物体生成;同时设计焦点区域增强数据管道(FRE-DP),强化生成中物体密集区域,并结合标注修正机制解决遗漏、多余和错位生成问题。大量实验表明,该方法显著优于现有先进方法,且可稳定提升多种检测器的性能。代码已开源:https://github.com/Sirius-Li/UAVGen。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout-to-image generation approaches have proved effective in promoting detection accuracy by synthesizing labeled images based on diffusion models. However, they suffer from frequently producing artifacts, especially near layout boundaries of tiny objects, thus substantially limiting their performance. To address these issues, we propose UAVGen, a novel layout-to-image generation framework tailored for UAV-based object detection. Specifically, UAVGen designs a Visual Prototype Conditioned Diffusion Model (VPC-DM) that constructs representative instances for each class and integrates them into latent embeddings for high-fidelity object generation. Moreover, a Focal Region Enhanced Data Pipeline (FRE-DP) is introduced to emphasize object-concentrated foreground regions in synthesis, combined with a label refinement to correct missing, extra and misaligned generations. Extensive experimental results demonstrate that our method significantly outperforms state-of-the-art approaches, and consistently promotes accuracy when integrated with distinct detectors. The source code is available at https://github.com/Sirius-Li/UAVGen.

无人机检测图像生成扩散模型小物体检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。