arXiv:2603.00163cs.CVcs.LG2026-03

针对白板笔画分割中极端类别不平衡问题,提出新评估协议以揭示模型真实性能。

A Boundary-Metric Evaluation Protocol for Whiteboard Stroke Segmentation Under Extreme Imbalance

  • 引入边界感知指标和薄笔画子集分析,打破传统指标误导
  • 重加权损失使F1提升20+点,边界精度同步改善
  • 揭示均值与最差表现的权衡,适合关注鲁棒性的研究者

白板笔画二值分割受极端类别不平衡影响,笔画像素平均仅占图像1.79%,薄笔画子集更低至1.14%±0.41%。标准区域指标(如F1、IoU)因背景主导而掩盖薄笔画失败。本文提出联合评估协议,包含区域与边界指标(BF1、B-IoU)、核心/薄笔画子集公平性分析、多轮训练下的每图鲁棒性统计(中位数、四分位距、最差情况)及非参数显著性检验。在DeepLabV3-MobileNetV3模型上,对五种损失函数(交叉熵、焦点损失、Dice、Dice+焦点、Tversky)各训练三次,于12张独立测试图像上评估。基于重叠的损失使F1达0.663,较交叉熵(0.438)提升超20点(p<0.001)。边界指标确认轮廓精度同步提升。原生分辨率下自适应阈值与Sauvola二值化平均F1达0.787,但最差情况降至0.452,显著低于Tversky的0.565,暴露一致性-准确性权衡:经典方法在均值上领先,学习模型更具最差情况可靠性。训练分辨率加倍进一步使F1提升12.7点。

原文摘要 · Abstract (English)

The binary segmentation of whiteboard strokes is hindered by extreme class imbalance, caused by stroke pixels that constitute only $1.79%$ of the image on average, and in addition, the thin-stroke subset averages $1.14% \pm 0.41%$ in the foreground. Standard region metrics (F1, IoU) can mask thin-stroke failures because the vast majority of the background dominates the score. In contrast, adding boundary-aware metrics and a thin-subset equity analysis changes how loss functions rank and exposes hidden trade-offs. We contribute an evaluation protocol that jointly examines region metrics, boundary metrics (BF1, B-IoU), a core/thin-subset equity analysis, and per-image robustness statistics (median, IQR, worst-case) under seeded, multi-run training with non-parametric significance testing. Five losses -- cross-entropy, focal, Dice, Dice+focal, and Tversky -- are trained three times each on a DeepLabV3-MobileNetV3 model and evaluated on 12 held-out images split into core and thin subsets. Overlap-based losses improve F1 by more than 20 points over cross-entropy ($0.663$ vs $0.438$, $p < 0.001$). In addition, the boundary metrics confirm that the gain extends to the precision of the contour. Adaptive thresholding and Sauvola binarization at native resolution achieve a higher mean F1 ($0.787$ for Sauvola) but with substantially worse worst-case performance (F1 $= 0.452$ vs $0.565$ for Tversky), exposing a consistency-accuracy trade-off: classical baselines lead on mean F1 while the learned model delivers higher worst-case reliability. Doubling training resolution further increases F1 by 12.7 points.

图像分割不平衡数据边界评估鲁棒性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。