arXiv:2609.07982cs.LG2026-09

研究空间采样偏差如何破坏半监督学习,发现性能会突然崩溃而非渐变。

Semi-Supervised Learning under Spatially Biased Sampling

论文配图:Semi-Supervised Learning under Spatially Biased Sampling
图 1 · 摘自论文原文
  • 通过控制参数分离空间偏差、自相关和非平稳性,构建可调合成数据框架
  • 当分布不匹配严重时,模型性能在0.71至0.77间出现阈值式骤降
  • 提出局部核加权发散度指标,更稳定检测空间分布差异,适合实际部署

标准半监督学习(SSL)通常假设标注与未标注数据具有相同的边缘分布。这一假设常被空间偏差采样机制打破,当标签在空间上存在选择性采集时尤为如此。本文将边缘分布不匹配、空间自相关和空间非平稳性视为三种独立机制,分别通过标注采样集中度参数、空间长度尺度和非平稳强度参数进行调控,探究分布不匹配如何削弱SSL效果,集群与流形假设是否仍成立,以及失败模式如何诊断。结合受控合成框架与PovertyMap-WILDS、加州住房、社会经济及美国空气质量监测数据集,系统调节不匹配程度并考虑空间自相关与非平稳性。通过分段回归变点分析发现,在合成生成器中,当分布不匹配足够严重时,SSL性能并非逐渐下降,而是在0.71至0.77之间呈现阈值式崩溃。进一步证明,空间非平稳性会独立于边缘不匹配导致性能损失,且模型在标签覆盖区域外变得愈发过度自信。为支持实际应用,评估多种分布差异度量作为可靠性指示器,并引入核加权局部发散度指标,相较于朴素局部方法提供更稳定的空间不匹配估计。这些发现为半监督学习中引入空间未标注数据的风险提供了实证依据与诊断工具。

原文摘要 · Abstract (English)

Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as three distinct mechanisms, varied independently via a labelled-sampling concentration parameter, a spatial length scale, and a non-stationarity strength parameter, and ask how mismatch degrades SSL, whether the cluster and manifold assumptions survive it, and how the resulting failure can be diagnosed. Using a controlled synthetic framework alongside PovertyMap-WILDS, California housing, socio-economic and US air quality monitoring datasets, we systematically vary the degree of mismatch while accounting for spatial autocorrelation and non-stationarity. Through a series of analyses including a segmented-regression changepoint, we show that in the synthetic generator, SSL performance does not degrade gradually but instead exhibits a threshold-like breakdown between approximately 0.71 and 0.77 once distribution mismatch becomes sufficiently severe. We further demonstrate that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available. To support practical deployment, we evaluate several distribution-divergence measures as indicators of reliability and introduce a kernel-weighted local divergence metric that provides a more stable estimate of spatial mismatch than a na\"ive localised approach. These findings provide empirical evidence and diagnostic tools for better documenting the risk of incorporating unlabelled spatial data into semi-supervised learning workflows.

半监督学习空间偏差分布不匹配诊断工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。