用视觉模型自动标注高分辨率3D军事场景,省时省力。
Open-Vocabulary High-Resolution 3D (OVHR3D) Data Segmentation and Annotation Framework
- 结合视觉大模型与3D网格渲染,实现跨模态标注
- 在军用3D数据上验证,显著降低人工标注成本
- 适合需要快速构建军事仿真环境的团队使用
在美军建模与仿真领域,高质量标注的3D数据对训练虚拟环境至关重要。传统3D语义和实例分割方法(如KpConv、RandLA、Mask3D)需依赖大规模标注数据集才能取得良好性能,但这类数据尤其在军事场景中稀缺。此前研究利用One World Terrain数据仓库的人工标注数据库,在IITSEC 2019和2021展示成果,扩展了训练数据。然而,针对特定任务的大规模3D数据收集与标注仍成本高昂、效率低下。为此,本研究旨在设计并开发一套高效全面的3D分割与标注框架,辅助3D数据标注。该框架融合Grounding DINO与Segment Anything Model,通过3D网格增强2D图像渲染效果,并开发了用户友好的交互界面,支持渲染图像与3D点云的直观可视化。
原文摘要 · Abstract (English)
In the domain of the U.S. Army modeling and simulation, the availability of high quality annotated 3D data is pivotal to creating virtual environments for training and simulations. Traditional methodologies for 3D semantic and instance segmentation, such as KpConv, RandLA, Mask3D, etc., are designed to train on extensive labeled datasets to obtain satisfactory performance in practical tasks. This requirement presents a significant challenge, given the inherent scarcity of manually annotated 3D datasets, particularly for the military use cases. Recognizing this gap, our previous research leverages the One World Terrain data repository manually annotated databases, as showcased at IITSEC 2019 and 2021, to enrich the training dataset for deep learning models. However, collecting and annotating large scale 3D data for specific tasks remains costly and inefficient. To this end, the objective of this research is to design and develop a comprehensive and efficient framework for 3D segmentation tasks to assist in 3D data annotation. This framework integrates Grounding DINO and Segment anything Model, augmented by an enhancement in 2D image rendering via 3D mesh. Furthermore, the authors have also developed a user friendly interface that facilitates the 3D annotation process, offering intuitive visualization of rendered images and the 3D point cloud.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。