arXiv:2608.09101cs.CVcs.LG2026-08

提出无需参考标注的遥感语义分割质量评估方法,直接检验掩码与图像证据的一致性。

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

论文配图:Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation
图 1 · 摘自论文原文
  • 通过对比反事实图像视图,让冻结的视觉语言模型判断掩码内是否集中了类别证据
  • 在10731个图像-类别对上发现:建筑类掩码62-85%更符合图像证据,而土地覆盖类则相反
  • 可识别标注偏差,适合用于提升遥感数据质量与模型跨域迁移能力

语义分割模型依赖人工绘制的掩码进行训练和评估,但遥感标注常存在粗糙、不完整或错位问题;高重叠率可能反映的是与错误标注的一致性,而非对图像的真实忠实度,造成评估困境。本文提出无需参考、无需训练的对比掩码保真度(CMF)度量方法,将候选掩码与图像证据直接对比。CMF通过保留或擦除掩码的反事实视图,让一个冻结的视觉-语言判别器判断某类证据是否集中在掩码内部且外部缺失。我们在受控的掩码破坏实验中验证了该方法,并在十项遥感基准上审计了10,731个图像-类别对,使用基于SegEarth-OV3构建的训练无关开放词汇探测器Seg-Probe生成的候选掩码。结果揭示系统性、类别依赖的标注失真:建筑物、道路、车辆等人造类别在62%-85%的样本中,候选掩码比人工标注更符合图像证据;而模糊的地表覆盖类则相反。在盲评三名标注者共识时,CMF在81%的样本中匹配专家判断,优于仅保留掩码、模型置信度及训练标签质量基线。最终,采用保守的逐类仲裁生成的监督信号,在跨域迁移任务上优于原始标注和匹配替换控制组,使CMF成为可扩展的真值审计工具,而非默认其绝对正确。

原文摘要 · Abstract (English)

Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.

遥感分割标注质量无参考评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。