arXiv:2606.09162cs.CV2026-06被引 3

用几何门控解决无人机视频语义分割的时序不一致问题

Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation

论文配图:Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation
图 1 · 摘自论文原文
  • 基于RANSAC内点率设计无参数几何门控,选择性使用单应变换或光流
  • 在合成数据集上提升4.24%-4.91% mIoU,时序一致性从62%提升至92%
  • 适合关注低空无人机视频分割稳定性的研究者和工程应用

低空无人机视频语义分割需保证时序一致性,但密集光流会在占主导地位的平面区域引入结构化噪声。本文提出一种零参数几何门控,通过在16×16空间网格上计算RANSAC单应变换内点率,决定每个区域采用单应变换或光流进行形变,再经语义相似性传播融合。该门控无需可学习参数,仅依赖中值阈值的二元决策,整体仅增加211K可训练参数(即语义相似性传播层)。在合成数据集UAVid上,该方法在两种架构(SegFormer-b2与Hiera-S+UPerNet)下相较基线模型提升4.24–4.91% mIoU。机制分析显示,平面区域的光流残差具有空间自相关性(莫兰指数0.32,p<0.001),可预测边界不稳定(斯皮尔曼ρ=0.66),且刚性化处理使有效区域内时序一致性从62%恢复至92%(+29.5个百分点)。

原文摘要 · Abstract (English)

Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the planar regions that dominate aerial imagery. We propose a zero-parameter geometric gate that uses RANSAC homography inlier ratios on a $16\times16$ spatial grid to route each region to either homography or optical flow warp before fusion via Semantic Similarity Propagation. The gate requires no learned parameters -- only a median-threshold binary decision on RANSAC statistics -- adding only 211K trainable parameters (the SSP fusion layer) to a frozen backbone. On synthetic UAVid, the method achieves +4.24--4.91\% mIoU improvement over base models across two architectures (SegFormer-b2 and Hiera-S+UPerNet). Mechanism diagnostics reveal that flow residuals in planar regions are spatially autocorrelated (Moran's I = 0.32, $p < 0.001$), predict boundary instability (Spearman $ρ= 0.66$), and that rigidification recovers temporal consistency from 62\% to 92\% (+29.5pp) in homography-valid regions.

语义分割无人机视频时序稳定几何门控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。