arXiv:2605.13755cs.CV2026-05

用生成模型批量打造不同外观的3D行人,提升自动驾驶感知鲁棒性。

Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception

论文配图:Generative Texture Diversification of 3D Pedestrians for Robust Autonomous Driving Perception
图 1 · 摘自论文原文
  • 基于StyleGAN2合成多样面部纹理并自动映射到3D网格
  • 合成数据使2D检测准确率提升12.3%,但3D检测受几何差异影响大
  • 适合做自动驾驶仿真数据增强与跨域训练策略研究者

近年来,自动驾驶对高质量数据的需求激增,用于训练安全关键场景下的2D和3D感知模型。真实世界数据集难以满足需求,因要求持续演进且大规模标注成本高、耗时长,合成数据成为可扩展、可控的替代方案。行人检测是自动驾驶中最关键的任务之一。本文提出一种简单有效的3D行人资产多样性扩展方法:从单一3D基础资产出发,利用StyleGAN2合成多样的面部纹理与身份级外观变化,并自动映射至3D网格,实现无需为每实例设计新几何的外观级多样化。使用这些资产构建合成数据集,研究了混合真实与合成数据对基于RGB的目标检测影响;通过互补实验,分析了点云感知中由几何驱动的分布偏移。结果表明,可控的合成多样化提升了2D检测鲁棒性,但揭示了3D感知模型对几何域差距的高度敏感性。整体上,本工作展示了生成式AI如何通过受控面部纹理合成,实现可扩展、可模拟的行人多样化,同时阐明了跨域训练策略在自动驾驶流程中的优势与局限。

原文摘要 · Abstract (English)

In recent years, autonomous driving has significantly in creased the demand for high-quality data to train 2D and 3D perception models for safety-critical scenarios. Real world datasets struggle to meet this demand as require ments continuously evolve and large-scale annotated data collection remains costly and time-consuming making syn thetic data a scalable, practical and controllable alterna tive. Pedestrian detection is among the most safety-critical tasks in autonomous driving. In this paper, we propose a simple yet effective method for scaling variability in 3D pedestrian assets for synthetic scene generation. Starting from a single 3D base asset, we generate multiple distinct pedestrian instances by synthesizing diverse facial textures and identity-level appearance variations using StyleGAN2 and automatically mapping them onto 3D meshes. This ap proach enables scalable appearance-level asset diversifica tion without requiring the design of new geometries for each instance. Using the assets, we construct synthetic datasets and study the impact of mixing real and synthetic data for RGB-based object detection. Through complementary ex periments, we analyze geometry-driven distribution shifts in point cloud perception for 3D object detection. Our findings demonstrate that controlled synthetic diversifica tion improves robustness in 2D detection while revealing the sensitivity of 3D perception models to geometric domain gaps. Overall, this work highlights how generative AI en ables scalable, simulation-ready pedestrian diversification through controlled facial texture synthesis, along with the benefits and limitations of cross-domain training strategies in autonomous driving pipelines.

3D行人生成模型自动驾驶数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。