arXiv:2511.08615cs.CVcs.IT2025-11

构建多无人机多视角数据集与追踪框架,解决动态拍摄中的遮挡难题。

A Multi-Drone Multi-View Dataset and Deep Learning Framework for Pedestrian Detection and Tracking

  • 基于八架无人机实时变位拍摄,实现鸟瞰图特征融合追踪
  • 复杂场景下检测与追踪准确率约90%,轨迹跟踪率达80%
  • 支持相机失效时平稳降级,适合真实安防部署

多无人机监控系统在行人追踪中具备更广覆盖和更强鲁棒性,但现有方法难以应对动态相机位姿与复杂遮挡。本文提出MATRIX(多空中复杂环境追踪)数据集,包含八架无人机同步拍摄的连续运动视频,以及一个新型深度学习框架。相比静态摄像机或有限无人机覆盖的数据集,MATRIX在城市环境中设置40名行人与显著建筑遮挡,构成高挑战场景。所提框架通过实时相机标定、基于特征的图像配准及鸟瞰图(BEV)表示下的多视角特征融合,应对动态拍摄挑战。实验表明,在无遮挡简化环境下,静态摄像头方法可保持90%以上精度;但在复杂环境中性能显著下降。本文方法在复杂条件下仍维持约90%的检测与追踪准确率,并成功追踪约80%的行进轨迹。迁移学习实验显示,预训练模型表现远优于从零训练,且系统相机断联实验验证了平滑性能退化,具备实际部署鲁棒性。MATRIX数据集与框架为动态多视角监控系统提供了关键基准。

原文摘要 · Abstract (English)

Multi-drone surveillance systems offer enhanced coverage and robustness for pedestrian tracking, yet existing approaches struggle with dynamic camera positions and complex occlusions. This paper introduces MATRIX (Multi-Aerial TRacking In compleX environments), a comprehensive dataset featuring synchronized footage from eight drones with continuously changing positions, and a novel deep learning framework for multi-view detection and tracking. Unlike existing datasets that rely on static cameras or limited drone coverage, MATRIX provides a challenging scenario with 40 pedestrians and a significant architectural obstruction in an urban environment. Our framework addresses the unique challenges of dynamic drone-based surveillance through real-time camera calibration, feature-based image registration, and multi-view feature fusion in bird's-eye-view (BEV) representation. Experimental results demonstrate that while static camera methods maintain over 90\% detection and tracking precision and accuracy metrics in a simplified MATRIX environment without an obstruction, 10 pedestrians and a much smaller observational area, their performance significantly degrades in the complex environment. Our proposed approach maintains robust performance with $\sim$90\% detection and tracking accuracy, as well as successfully tracks $\sim$80\% of trajectories under challenging conditions. Transfer learning experiments reveal strong generalization capabilities, with the pretrained model achieving much higher detection and tracking accuracy performance compared to training the model from scratch. Additionally, systematic camera dropout experiments reveal graceful performance degradation, demonstrating practical robustness for real-world deployments where camera failures may occur. The MATRIX dataset and framework provide essential benchmarks for advancing dynamic multi-view surveillance systems.

多无人机行人追踪鸟瞰图数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。