arXiv:2603.24294cs.CV2026-03

用视觉先验生成更真实的稀有3D物体,提升长尾检测效果

VERIA: Verification-Centric Multimodal Instance Augmentation for Long-Tailed 3D Object Detection

  • 基于图像生成同步的RGB-LiDAR实例,融合现成大模型
  • 在nuScenes和Lyft上显著提升稀有类别检测性能
  • 通过多阶段验证确保生成数据真实且多样,适合长尾检测研究

驾驶数据集中的长尾分布给3D感知带来挑战,稀有类别虽内部差异大,但样本覆盖不足。现有基于复制粘贴或资产库的实例增强方法在细粒度多样性与场景上下文放置上受限。我们提出VERIA,一种以图像为主的多模态增强框架,利用现成基础模型合成同步的RGB-LiDAR实例,并通过顺序语义与几何验证进行筛选。该验证驱动设计能选出更符合真实LiDAR统计特征、同时覆盖更广类内变异的实例。分阶段产出分解提供管道可靠性日志诊断。在nuScenes和Lyft数据集上,VERIA在仅使用LiDAR及多模态设置下均提升了稀有类别的3D物体检测性能。代码已公开于https://sgvr.kaist.ac.kr/VERIA/。

原文摘要 · Abstract (English)

Long-tail distributions in driving datasets pose a fundamental challenge for 3D perception, as rare classes exhibit substantial intra-class diversity yet available samples cover this variation space only sparsely. Existing instance augmentation methods based on copy-paste or asset libraries improve rare-class exposure but are often limited in fine-grained diversity and scene-context placement. We propose VERIA, an image-first multimodal augmentation framework that synthesizes synchronized RGB--LiDAR instances using off-the-shelf foundation models and curates them with sequential semantic and geometric verification. This verification-centric design tends to select instances that better match real LiDAR statistics while spanning a wider range of intra-class variation. Stage-wise yield decomposition provides a log-based diagnostic of pipeline reliability. On nuScenes and Lyft, VERIA improves rare-class 3D object detection in both LiDAR-only and multimodal settings. Our code is available at https://sgvr.kaist.ac.kr/VERIA/.

3D检测长尾问题数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。