arXiv:2504.17076cs.CV2025-04被引 2

提出场景感知布局模型,让生成物体更真实地分布于汽车检测图像中。

Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection

  • 基于场景概率建模,预测物体在图像中的合理位置。
  • 在真实布局上生成物体,使检测模型提升1.4 mAP,优于此前最佳方案。
  • 适合自动驾驶视觉任务的数据增强,尤其对布局敏感的检测与分割任务。

生成式图像模型在视觉任务数据增强中日益普及。在汽车目标检测中,现有方法通常追求图像外观的逼真性(如用生成物体替换真实物体),或最大化画面多样性(如将大量生成物体粘贴到背景上)。但这些方法大多忽视了物体在场景中的合理位置:要么重复使用原有布局,要么随机放置,完全脱离真实感。本文主张,最优的数据增强应包含合理的场景布局生成。为此,我们提出一种场景感知的概率化位置建模方法,可预测新物体在真实场景中可合理放置的位置。随后,利用生成模型在这些位置进行图像修复(inpainting),获得更强的增强效果。我们在两个汽车目标检测任务上实现了生成式数据增强的新基准,性能提升最高达2.8倍(+1.4 vs. +0.5 mAP),显著优于现有方法;同时在实例分割任务上也取得明显改进。

原文摘要 · Abstract (English)

Generative image models are increasingly being used for training data augmentation in vision tasks. In the context of automotive object detection, methods usually focus on producing augmented frames that look as realistic as possible, for example by replacing real objects with generated ones. Others try to maximize the diversity of augmented frames, for example by pasting lots of generated objects onto existing backgrounds. Both perspectives pay little attention to the locations of objects in the scene. Frame layouts are either reused with little or no modification, or they are random and disregard realism entirely. In this work, we argue that optimal data augmentation should also include realistic augmentation of layouts. We introduce a scene-aware probabilistic location model that predicts where new objects can realistically be placed in an existing scene. By then inpainting objects in these locations with a generative model, we obtain much stronger augmentation performance than existing approaches. We set a new state of the art for generative data augmentation on two automotive object detection tasks, achieving up to $2.8\times$ higher gains than the best competing approach ($+1.4$ vs. $+0.5$ mAP boost). We also demonstrate significant improvements for instance segmentation.

数据增强目标检测生成模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。