通过修正难样本语义表示,提升医学图像分割的准确性。
SHTA: Semantic Hard Token Correction and Center Alignment for Semi-Supervised Medical Image Segmentation

- 引入语义修正与中心对齐机制,优化难区域特征表达
- 在多个框架上实现分割精度和弱器官恢复显著提升
- 轻量级设计,仅增加训练开销,无推理成本
近期半监督医学图像分割方法通过预测一致性、伪标签监督和难区域监督取得了显著进展。然而,这些方法主要提升监督质量,未显式强制难区域学习到语义一致的表示。因此,即使预测层面监督增强,难区域仍常出现语义分配不稳定,阻碍进一步性能提升。为此,我们提出SHTA(语义难令牌修正与中心对齐),一个轻量级训练时语义表示分支。SHTA不引入额外预测监督,而是通过语义分配、难令牌精炼和语义中心对齐,改进难区域的语义一致性,同时保留原有预测路径,且无额外推理开销。我们将SHTA集成至GA-CPS、CPS、URPC和MagicNet等代表性框架,在Synapse和AMOS数据集上评估。结果表明,SHTA在各框架中均带来一致提升,尤其在分割精度、弱器官恢复和语义模糊性降低方面效果显著,仅增加训练时间开销。代码已公开于https://anonymous.4open.science/r/release_SHTA-42D5/。
原文摘要 · Abstract (English)
Recent advances in semi-supervised medical image segmentation have achieved remarkable performance through prediction consistency, pseudo-label supervision, and hard-region supervision. However, these methods primarily improve supervision quality rather than explicitly enforcing semantic consistency in the learned representations of hard regions. Consequently, even under increasingly stronger prediction-level supervision, difficult regions exhibiting unstable semantic assignment often fail to establish semantically consistent representations during training, thereby limiting further segmentation improvement. To address this issue, we propose SHTA (Semantic Hard Token Correction and Center Alignment), a lightweight training-time semantic representation branch. Instead of introducing additional prediction supervision, SHTA refines intermediate semantic representations through Semantic Assignment, Hard Token Refinement, and Semantic Center Alignment, thereby improving semantic consistency in hard regions while preserving the original prediction pathway and introducing no additional inference cost. We integrate SHTA into representative semi-supervised segmentation frameworks, including GA-CPS, CPS, URPC, and MagicNet, and conduct evaluations on the Synapse and AMOS datasets. Experimental results demonstrate that SHTA delivers consistent paired improvements across frameworks, with especially clear gains in segmentation accuracy, weak-organ recovery, and semantic ambiguity reduction, while incurring only training-time overhead. The code is available at https://anonymous.4open.science/r/release_SHTA-42D5/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。