无需先验信息的无人机3D姿态估计算法,支持实时应用且泛化能力强。
DroneKey++: A Size Prior-free Method and New Benchmark for Drone 3D Pose Estimation from Sequential Images
- 不依赖尺寸或模型数据,通过关键点检测与几何推理联合估计姿态。
- 旋转误差17.34度,平移误差0.135米,推理速度达414帧/秒(GPU)。
- 新构建6DroneSyn数据集,覆盖7种机型、88种背景,适合真实场景验证。
精确的无人机3D姿态估计对安防系统至关重要。然而现有方法通常依赖物理尺寸或3D模型等先验信息,且现有数据集规模小、仅限单一机型、采集环境受限,难以可靠评估泛化能力。我们提出DroneKey++,一种无需先验信息的框架,可联合完成关键点检测、无人机分类与3D姿态估计。该框架采用关键点编码器实现关键点检测与分类同步处理,通过基于射线的几何推理和类别嵌入估计3D姿态。为解决数据集局限,我们构建了6DroneSyn,一个大规模合成基准,包含超过5万张图像,涵盖7种无人机型号和88种室外背景,采用360°全景合成生成。实验表明,DroneKey++在旋转估计上达到17.34度(MAE)和17.1度(MedAE),平移估计为0.135米(MAE)和0.242米(MedAE),CPU推理速度19.25 FPS,GPU达414.07 FPS,展现出强跨机型泛化能力与实时应用潜力。数据集已公开。
原文摘要 · Abstract (English)
Accurate 3D pose estimation of drones is essential for security and surveillance systems. However, existing methods often rely on prior drone information such as physical sizes or 3D meshes. At the same time, current datasets are small-scale, limited to single models, and collected under constrained environments, which makes reliable validation of generalization difficult. We present DroneKey++, a prior-free framework that jointly performs keypoint detection, drone classification, and 3D pose estimation. The framework employs a keypoint encoder for simultaneous keypoint detection and classification, and a pose decoder that estimates 3D pose using ray-based geometric reasoning and class embeddings. To address dataset limitations, we construct 6DroneSyn, a large-scale synthetic benchmark with over 50K images covering 7 drone models and 88 outdoor backgrounds, generated using 360-degree panoramic synthesis. Experiments show that DroneKey++ achieves MAE 17.34 deg and MedAE 17.1 deg for rotation, MAE 0.135 m and MedAE 0.242 m for translation, with inference speeds of 19.25 FPS (CPU) and 414.07 FPS (GPU), demonstrating both strong generalization across drone models and suitability for real-time applications. The dataset is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。