arXiv:2409.18300cs.CVcs.AI2024-09被引 1

FALCON让无人机视频识别更准更快,专注关键动作区域。

FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition

  • 用目标感知掩码自编码+未来内容重建,聚焦动作相关区域。
  • 在NEC-Drone上提升2.9%,UAV-Human上提升5.8%。
  • 无需推理时预处理,速度比传统方法快2到5倍。

我们提出FALCON,一种统一的自监督视频预训练方法,用于从原始RGB航拍视频中进行无人机动作识别,推理时无需额外预处理。无人机视频存在严重空间失衡:大而杂乱的背景占据视野,导致基于重建的预训练过度关注无信息区域,忽视动作相关的人员或物体线索。FALCON通过引入目标感知掩码自编码与目标中心双时域未来重建来解决此问题。仅在预训练阶段使用检测结果构建目标度先验,实现(i)遮蔽时平衡标记可见性,(ii)将重建监督集中于动作相关区域,防止学习被背景外观主导。为进一步促进时序动态学习,我们在目标中心监督区域内重建短时和长时未来内容,注入对噪声航拍上下文鲁棒的前瞻时序监督。在多个无人机基准上,使用ViT-B骨干网络,FALCON在NEC-Drone上提升顶1准确率2.9%,在UAV-Human上提升5.8%,且推理速度比依赖复杂测试增强的监督方法快2×–5×。

原文摘要 · Abstract (English)

We introduce FALCON, a unified self-supervised video pretraining approach for UAV action recognition from raw RGB aerial footage, requiring no additional preprocessing at inference. UAV videos exhibit severe spatial imbalance: large, cluttered backgrounds dominate the field of view, causing reconstruction-based pretraining to waste capacity on uninformative regions and under-learn action-relevant human/object cues. FALCON addresses this by integrating object-aware masked autoencoding with object-centric dual-horizon future reconstruction. Using detections only during pretraining, we construct objectness priors that (i) enforce balanced token visibility during masking and (ii) concentrate reconstruction supervision on action-relevant regions, preventing learning from being dominated by background appearance. To promote temporal dynamics learning, we further reconstruct short- and long-horizon future content within an object-centric supervision region, injecting anticipatory temporal supervision that is robust to noisy aerial context. Across UAV benchmarks, FALCON improves top-1 accuracy by 2.9\% on NEC-Drone and 5.8\% on UAV-Human with a ViT-B backbone, while achieving 2$\times$--5$\times$ faster inference than supervised approaches that rely on heavy test-time augmentation.

无人机识别自监督学习目标中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。