解决无人机多光谱视频分割中的模态差异与时间不一致问题
MSTF-Net: A UAV-Oriented Multi-Spectral Video Segmentation Method via Modality-Robust, Scale-Adaptive, and Consistent Fusion

- 通过跨模态注意力与残差引导过滤,实现RGB与热成像的鲁棒融合
- 在MVSeg和CART数据集上分别达到56.42%和51.80% mIoU,性能领先
- 适合处理小目标、遮挡及模态退化等复杂场景的无人机视觉应用
多光谱视频分割对无人机在城市规划、土地利用监测、交通监控和人群估计等应用中的鲁棒场景理解至关重要。尽管RGB与热成像的融合能提供互补信息,应对光照与能见度变化,但仍面临两大挑战:(1) 模态融合困境,源于RGB与热成像特征显著差异导致互补线索被掩盖;(2) 时间变化,由无人机平台快速运动和视角变化引发帧间外观不一致与错位。为此,本文提出MSTF-Net,一种面向无人机的模态鲁棒、尺度自适应融合框架,有效建模跨模态融合与时间一致性。模态空间互补抑制增强(MSCSE)模块通过跨模态注意力生成统一实例查询,并利用残差引导差异过滤与一致性约束抑制模态特异性噪声。为建模时间动态,多尺度时序跨模态语义一致性(MTCSC)模块根据帧距自适应调整时序感受野,捕捉时间上的粗粒度全局上下文与细粒度局部结构。在公开的RGB-T数据集上的大量消融实验表明,MSTF-Net在复杂条件下(如小目标、遮挡、模态退化)均取得当前最优分割性能,分别在MVSeg数据集上达到56.42% mIoU,CART数据集上达到51.80% mIoU。
原文摘要 · Abstract (English)
Multi-spectral video segmentation is essential for robust scene understanding in unmanned aerial vehicle (UAV) applications such as city planning, land use monitoring, traffic monitoring, and crowd estimation. While the fusion of RGB and thermal modalities offers complementary information for perception under varying lighting and visibility conditions, two fundamental challenges remain: (1) the modal fusion dilemma, arising from significant discrepancies between RGB and thermal features that obscure complementary cues, and (2) temporal variation, induced by rapid motion and viewpoint changes on UAV platforms, which leads to appearance inconsistency and misalignment across frames. To address these issues, this study proposed MSTF-Net, a modality-robust scale-adaptive fusion framework for multi-spectral video segmentation that effectively models cross-modal fusion and temporal consistency. The Modality Spatial Complementary Suppression and Enhancement (MSCSE) module generates unified instance queries via cross-modal attention and suppresses modality-specific noise using residual-guided discrepancy filtering and consistency constraints. To model temporal dynamics, the Multi-scale Temporal Cross-modality Semantic Consistency (MTCSC) module adaptively adjusts the temporal receptive field based on frame distance, capturing both coarse global context and fine local structure across time. Extensive ablation experiments on public RGB-T datasets demonstrate that MSTF-Net achieves state-of-the-art segmentation performance, especially under challenging conditions such as small targets, occlusion, and modality degradation. Specifically, reached 56.42\% mIoU on the MVSeg dataset and 51.80\% mIoU on the CART dataset, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。