提升越野语义分割边界精度,减少噪声干扰
Terrain-Enhanced Resolution-aware Refinement Attention for Off-Road Segmentation
- 低分辨率瓶颈+门控跨注意力注入细节,高效保持边缘清晰
- 仅对不确定区域进行稀疏精修,计算少且抗噪强
- 适合弱监督场景,尤其对罕见类别分割效果好
越野语义分割面临边界模糊、稀疏标注和普遍标签噪声等问题。仅在低分辨率融合会模糊边缘并传播局部错误,而维持高分辨率路径或重复高分辨率融合则成本高且易受噪声影响。本文提出一种分辨率感知的令牌解码器,在不完美监督下兼顾全局语义、局部一致性和边界保真度。大部分计算在低分辨率瓶颈中完成;通过门控跨注意力注入细粒度特征,仅对稀疏且不确定性高的像素进行精修。各组件协同设计:带轻量空洞深度卷积的自注意力恢复局部一致性;门控跨注意力融合标准高分辨率编码流特征,不放大噪声;类别感知点精修以极小开销修正残余歧义。训练时引入边界带一致性正则化,鼓励标注边缘周围薄邻域内预测一致,无推理开销。整体表现具有竞争力,跨过渡区稳定性显著提升。
原文摘要 · Abstract (English)
Off-road semantic segmentation suffers from thick, inconsistent boundaries, sparse supervision for rare classes, and pervasive label noise. Designs that fuse only at low resolution blur edges and propagate local errors, whereas maintaining high-resolution pathways or repeating high-resolution fusions is costly and fragile to noise. We introduce a resolutionaware token decoder that balances global semantics, local consistency, and boundary fidelity under imperfect supervision. Most computation occurs at a low-resolution bottleneck; a gated cross-attention injects fine-scale detail, and only a sparse, uncertainty-selected set of pixels is refined. The components are co-designed and tightly integrated: global self-attention with lightweight dilated depthwise refinement restores local coherence; a gated cross-attention integrates fine-scale features from a standard high-resolution encoder stream without amplifying noise; and a class-aware point refinement corrects residual ambiguities with negligible overhead. During training, we add a boundary-band consistency regularizer that encourages coherent predictions in a thin neighborhood around annotated edges, with no inference-time cost. Overall, the results indicate competitive performance and improved stability across transitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。