让医学影像分割模型同时捕捉专家差异与设备噪声,生成更可信的个性化结果。
Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

- 通过频域提示和动态特征调制,分离设备噪声与真实标注差异。
- 在LIDC-IDRI和NPC-170数据集上,Dice分数提升且不确定性更符合临床判断。
- 适合需要处理多专家标注、关注分割置信度的临床场景。
多评分者医学图像分割旨在捕捉临床解读中的固有模糊性,即不同专家和成像设备间诊断边界存在差异。现有方法常将多样性简化为共识标签或视评分者差异为噪声,导致模型过度自信且校准不足。本文提出一种谐化概率框架,通过自适应特征调节与频域个性化,分离采集伪影与真实标注变异。轻量级谐化网络隐式建模扫描仪特异性伪影,并动态调制特征以标准化潜在表示,确保不确定性反映解剖结构而非噪声。引入新型高频提示模块,在频域编码评分者依赖的边界精度与纹理敏感性,自适应调制谐化特征以生成个性化且解剖一致的分割结果。此外,基于广义能量距离的正则化使生成分布匹配实际标注变异,在专家分歧处保留多样性,共识区域趋于一致。在LIDC-IDRI和NPC-170数据集上的实验表明,该方法在聚合与个体化分割上均达到最先进性能,显著降低GED值并提升Dice分数,尤其在噪声病例中表现突出。除准确性外,模型展现出临床意义的不确定性:一致性区域置信度上升,模糊区域下降,支持其作为可靠可解释的多专家临床工作流工具。
原文摘要 · Abstract (English)
Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and imaging devices. Existing approaches often reduce this diversity to consensus labels or treat rater differences as noise, resulting in overconfident and poorly calibrated models. We propose a harmonized probabilistic framework that disentangles acquisition artifacts from genuine annotator variability through adaptive feature conditioning and frequency-domain personalization. A lightweight Harmonizer Network implicitly models scanner-specific artifacts and performs dynamic feature modulation to standardize latent representations, ensuring that uncertainty reflects anatomy rather than noise. To represent rater-specific styles, we introduce a novel High-Frequency Prompt Modules that operate in the spectral domain to encode annotator-dependent boundary precision and textural sensitivity. These prompts adaptively modulate harmonized features to produce personalized yet anatomically consistent segmentations. Furthermore, a Generalized Energy Distance based regularization aligns the generative distribution with empirical annotation variability, promoting diversity where experts disagree and consensus where they converge. Experiments on LIDC-IDRI and NPC-170 show SOTA aggregated and individualized segmentation, with notable GED reductions and improved Dice scores, especially on noisy cases. Beyond accuracy, the model exhibits clinically meaningful uncertainty. Confidence rises in agreement regions and declines in ambiguous areas, supporting its use as a reliable and interpretable tool for multi-expert clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。