arXiv:2606.29783cs.ROcs.AI2026-06被引 1

用自动标注生成10万张真实感图像,实现无人机无GPS追踪零样本迁移。

FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking

论文配图:FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking
图 1 · 摘自论文原文
  • 基于高斯点云模拟器自动标注,20分钟生成1万张带掩码与位姿标签的图像。
  • 零样本迁移下对三类物体实现96%-100%识别准确率,真实场景追踪成功率100%。
  • 适合做视觉导航与物理感知融合的无人机系统研究者参考。

视觉引导的空中追踪在无GPS环境中至关重要。可靠的追踪依赖大规模标注数据,但多数真实感数据集需大量人工标注且制作耗时。本文提出FalconTrack,一个统一的感知与追踪框架:(i) 利用可编辑的光栅化模拟器实现自动化标签生成;(ii) 结合多头感知与物理感知追踪,实现零样本从仿真到现实的迁移。FalconTrack在高斯点云模拟器中构建自动化标注流程,将目标高斯对象从短视频中分离,并与随机背景合成,生成约10,000张包含RGB、掩码、类别和6-DoF姿态标签的图像,耗时不足20分钟。基于该数据集,我们训练了分阶段学习并具有重投影一致性的多头感知模块,并将其输出与类别条件的动力学先验融合至扩展卡尔曼滤波器(EKF)中进行追踪。感知模型在三个几何形态各异的物体及两个环境上实现了96%-100%的零样本仿真到现实转移准确率,且在未见的仿真与真实场景中保持稳定表现。在真实硬件闭环视觉追踪中,机载系统运行频率约为25 Hz,五条轨迹跨越两个环境的F1-tenth与门控追踪任务均达到100%成功,而以掩码为中心的视觉基线在快速离屏场景下成功率下降至60%。

原文摘要 · Abstract (English)

Vision-based aerial tracking is critical in GPS-denied environments. Reliable perception for tracking depends on large-scale labeled data, yet most photorealistic datasets rely on heavy manual annotation and are time-consuming to produce. We present FalconTrack, a unified perception-and-tracking framework that (i) leverages a photorealistic editable simulator for automated label generation and (ii) combines multi-head perception with physics-aware tracking for zero-shot sim-to-real transfer. FalconTrack provides an automated labeling pipeline in a Gaussian Splatting simulator that isolates target Gaussians from short object videos and composites them with randomized backgrounds to generate RGB, mask, class, and 6-DoF pose labels, producing about 10k labeled images in under 20 minutes. Using this dataset, we train a multi-head perception module with staged learning and reprojection consistency, and fuse its outputs with class-conditioned dynamics priors in an EKF for tracking. Our perception model outperforms two baselines and reaches 96-100% class accuracy in zero-shot sim-to-real transfer on three geometrically diverse objects and two environments, while maintaining consistent performance in unseen simulated and real scenes. In real hardware closed-loop visual tracking, the onboard system runs at about 25 Hz and achieves 100% success in sim-to-real F1-tenth and gate tracking in five trajectories across two environments, while a mask-centered vision baseline drops to 60% success on F1-tenth during fast out-of-view scenarios.

无人机追踪自动标注零样本迁移物理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。