用3D高斯点云构建无人机视角的高保真数字孪生,增强动态目标检测训练
UAVTwin: Neural Digital Twins for UAVs using Gaussian Splatting
- 融合3D高斯点云与可控人形模型,从无人机视角合成复杂场景中的动态人物
- 相比现有方法提升1.23 dB PSNR,人检测任务mAP提高2.5%至13.7%
- 适合需要真实场景数据增强的无人机感知系统研发人员
我们提出UAVTwin,一种基于真实环境创建无人机数字孪生的方法,用于增强嵌入式无人飞行器(UAV)下游模型的训练数据。该方法聚焦于从无人机视角合成前景组件,如复杂背景中运动的多样化人类实例。通过结合3D高斯点云(3DGS)重建背景,以及可控制的合成人形模型(呈现多种外观与姿态),实现高保真渲染。据我们所知,UAVTwin是首个基于3DGS的无人机感知数字孪生方法。针对真实环境中多动态物体和显著外观变化带来的3DGS建模伪影问题,我们提出新型外观建模策略与掩码优化模块,有效提升3DGS训练效果。神经渲染质量方面,相较最新方法实现1.23 dB PSNR提升;在人检测任务中,数据增强使mAP提升2.5%至13.7%。
原文摘要 · Abstract (English)
We present UAVTwin, a method for creating digital twins from real-world environments and facilitating data augmentation for training downstream models embedded in unmanned aerial vehicles (UAVs). Specifically, our approach focuses on synthesizing foreground components, such as various human instances in motion within complex scene backgrounds, from UAV perspectives. This is achieved by integrating 3D Gaussian Splatting (3DGS) for reconstructing backgrounds along with controllable synthetic human models that display diverse appearances and actions in multiple poses. To the best of our knowledge, UAVTwin is the first approach for UAV-based perception that is capable of generating high-fidelity digital twins based on 3DGS. The proposed work significantly enhances downstream models through data augmentation for real-world environments with multiple dynamic objects and significant appearance variations-both of which typically introduce artifacts in 3DGS-based modeling. To tackle these challenges, we propose a novel appearance modeling strategy and a mask refinement module to enhance the training of 3D Gaussian Splatting. We demonstrate the high quality of neural rendering by achieving a 1.23 dB improvement in PSNR compared to recent methods. Furthermore, we validate the effectiveness of data augmentation by showing a 2.5% to 13.7% improvement in mAP for the human detection task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。