arXiv:2606.15112cs.CV2026-06

通过时间一致性学习提升卫星视频中细粒度目标检测精度

Learn Temporal Consistency For Robust Satellite Video Detector

论文配图:Learn Temporal Consistency For Robust Satellite Video Detector
图 1 · 摘自论文原文
  • 利用视频时序上下文,融合多帧特征实现目标一致表征
  • 在SAT-MTB数据集上达到47.7% mAP,较基线提升4.8%
  • 适配现有图像检测器,适用于卫星遥感中的细粒度识别

面向卫星视频中定向且细粒度目标的检测在遥感应用中具有重要意义。现有方法通常仅关注少数粗粒度移动目标类别,并使用水平边界框表示,难以提取卫星视频中目标的完整、准确和一致信息。本文提出基于时间一致性学习(TCL)的卫星视频目标检测框架,通过利用视频中丰富的时序上下文,有效检测定向与细粒度目标。该框架集成三个关键模块:时序与细粒度特征聚合(TFA)、结构编码(SE)和时间一致性约束(TCC)。TFA与TCC模块促进跨帧一致表征学习,而SE模块同时编码外观与结构信息,实现精准细粒度识别。在SAT-MTB基准数据集上的实验表明,TCL取得47.7% mAP的新最佳性能,较基线提升4.8%。此外,该框架可无缝兼容现有图像级检测器,显著提升检测精度。

原文摘要 · Abstract (English)

Satellite video object detection (SVOD) for oriented and fine-grained objects plays an important role in satellite applications. Most existing SVOD methods only focus on one or a few coarse-grained categories of moving objects and represent objects with horizontal bounding boxes. They have difficulty extracting complete, accurate, and consistent information about objects in whole satellite videos. In this paper, we propose a satellite video object detection framework based on Temporal Consistency Learning (TCL). TCL adeptly detects oriented and fine-grained objects by leveraging the rich temporal contexts within satellite videos. The framework integrates three key modules: temporal and fine-grained feature aggregation (TFA), structure encoding (SE), and temporal consistency constraint (TCC). TFA and TCC modules facilitate consistent representation learning across frames, while the SE module encodes both appearance and structural information for precise fine-grained recognition. Experimental results on the SAT-MTB benchmark dataset demonstrate TCL's superior performance, achieving a new state-of-the-art oriented and fine-grained detection accuracy of 47.7% mAP--a 4.8% improvement over the baseline. Furthermore, our TCL framework readily accommodates existing image-based detectors, leading to enhanced detection accuracies.

卫星视频目标检测时间一致性细粒度识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。