arXiv:2512.13411cs.CVcs.GR2025-12

用高斯点阵生成机器人视觉合成数据,自动标注且逼真。

Computer vision training dataset generation for robotic environments using Gaussian splatting

  • 用3D高斯点阵构建真实感环境,结合物理模拟排列物体。
  • 两阶段渲染加阴影图合成,显著提升图像真实感。
  • 自动生成像素级分割掩码,适合直接训练检测模型。

本文提出一种全新流程,用于生成大规模、高度逼真且自动标注的机器人环境计算机视觉训练数据集。该方法解决合成图像与真实世界之间的领域差距以及人工标注耗时两大难题。利用3D高斯点阵(3DGS)创建操作环境与物体的逼真三维表示,并在游戏引擎中通过物理模拟生成自然布局。采用新颖的两阶段渲染技术,将点阵的写实效果与代理网格生成的阴影图相结合,算法化融合后添加物理上合理的阴影与细微高光,大幅提升真实感。像素级分割掩码可自动生成,并直接适配如YOLO等目标检测模型。实验表明,将少量真实图像与大量合成数据混合训练,能取得最佳检测与分割性能,验证了该策略在高效构建鲁棒准确模型中的优越性。

原文摘要 · Abstract (English)

This paper introduces a novel pipeline for generating large-scale, highly realistic, and automatically labeled datasets for computer vision tasks in robotic environments. Our approach addresses the critical challenges of the domain gap between synthetic and real-world imagery and the time-consuming bottleneck of manual annotation. We leverage 3D Gaussian Splatting (3DGS) to create photorealistic representations of the operational environment and objects. These assets are then used in a game engine where physics simulations create natural arrangements. A novel, two-pass rendering technique combines the realism of splats with a shadow map generated from proxy meshes. This map is then algorithmically composited with the image to add both physically plausible shadows and subtle highlights, significantly enhancing realism. Pixel-perfect segmentation masks are generated automatically and formatted for direct use with object detection models like YOLO. Our experiments show that a hybrid training strategy, combining a small set of real images with a large volume of our synthetic data, yields the best detection and segmentation performance, confirming this as an optimal strategy for efficiently achieving robust and accurate models.

3D生成数据合成机器人视觉自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。