构建遥感建筑分割噪声数据集与评估基准,提升标注质量可靠性。
Data-Centric Benchmark for Label Noise Estimation and Ranking in Remote Sensing Binary Building Segmentation
- 设计可控扰动的遥感二值建筑分割数据集,支持噪声实验。
- 提出基于模型不确定性和一致性分析的噪声识别方法,优于现有基线。
- 适合遥感图像分割、数据质量评估与主动学习研究者使用。
高质量像素级标注对遥感图像语义分割至关重要,但因其人工标注成本高且耗时,常出现标注错误,严重影响现代分割模型的性能与鲁棒性。为应对这一挑战,本文提出一个全新的数据驱动型基准,并发布一个公开可用的二值建筑分割数据集。该任务作为典型且实用的测试平台,可实现对不同标注扰动的受控实验。同时,本文还引入两种技术,用于识别、量化并按噪声程度排序训练样本。所提方法结合模型不确定性、预测一致性和表示分析等互补策略,在多种实验设置中均显著优于已有基线。相关成果已开源:https://github.com/keillernogueira/label_noise_segmentation。
原文摘要 · Abstract (English)
High-quality pixel-level annotations are essential for the semantic segmentation of remote sensing imagery. However, such labels are expensive to obtain and often affected by noise due to the labor-intensive and time-consuming nature of pixel-wise annotation, which makes it challenging for human annotators to label every pixel accurately. Annotation errors can significantly degrade the performance and robustness of modern segmentation models, motivating the need for reliable mechanisms to identify and quantify noisy training samples. This paper introduces a novel data-centric benchmark, together with a new, publicly available binary building segmentation dataset. This specific task serves as a representative and practically relevant testbed that enables controlled experimentation with different annotation perturbations. Furthermore, we also introduce two techniques for identifying, quantifying, and ranking training samples according to their level of label noise in remote sensing semantic segmentation. Such proposed methods leverage complementary strategies based on model uncertainty, prediction consistency, and representation analysis, and consistently outperform established baselines across a range of experimental settings. The outcomes of this work are publicly available at https://github.com/keillernogueira/label_noise_segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。