医学影像编辑需评估整体语义漂移,而不仅是目标病灶得分。
Beyond Target Scores: Measuring Off-Target Drift in Diffusion-Based Medical Image Editing

- 构建胸部X光轨迹级评测基准CIB-Med-1,评估多种干扰因素下的编辑稳定性。
- 无约束编辑导致非目标区域语义漂移显著(中位数0.46,90%分位0.98)。
- 约束式扩散引导降低漂移至0.20和0.33,更贴近临床真实进展顺序。
扩散模型可实现视觉上合理的医学图像编辑,但传统评估仅关注目标分数提升,忽略了临床影像中目标病灶与共病、成像效应及选择偏倚的纠缠关系。模型可能通过改变相关非目标特征而非精准定位病理来‘伪装成功’。本文提出CIB-Med-1,一个用于胸部放射学中受控生物标志物编辑的轨迹级基准,通过校准的目标进展、反转率及14个临床相关干扰轴上的非目标语义漂移进行评估。该基准揭示了奖励欺骗问题:扩散编辑器在提升积液评分的同时,同时改变了肺实质、心胸轮廓、慢性病变或伪影相关特征。我们进一步提出一种约束扩散引导基线,在控制非目标变化范围内优化目标进展。在独立测试集上,该方法保持目标进展能力(ρ_trend=0.88 vs. 0.90),将中位数非目标漂移从0.46降至0.20,90%分位从0.98降至0.33。漂移幅度与目标-非目标关联性实证一致,表明语义不稳定性具有结构性而非偶然。盲测人类验证显示,放射科住院医师对意图进展排序的共识度显著提高(τ=0.61 vs. 0.29,Pix2Pix)。结果表明,医学图像编辑应以轨迹级语义控制为评价标准,而非终点分数最大化。
原文摘要 · Abstract (English)
Diffusion models can now edit medical images in visually plausible ways, but the standard evaluation question is too narrow: did the target score increase? In clinical imaging, target findings are entangled with co-morbidities, acquisition effects, and selection bias, so a model can appear successful by changing correlated non-target findings rather than isolating the intended pathology. We introduce CIB-Med-1, a trajectory-level benchmark for controlled biomarker editing in chest radiography. CIB-Med-1 evaluates directional pleural effusion editing through calibrated target progression, inversion rate, and off-target semantic drift over 14 clinically motivated nuisance axes. The benchmark exposes a reward-hacking failure mode in which diffusion editors increase effusion scores while simultaneously altering parenchymal, cardiomediastinal, pleural, chronic, or artifact-related findings. We further present a constrained diffusion guidance baseline that optimizes target progression subject to bounded off-target change. Across held-out radiographs, the constrained editor preserves target progression ($ρ_{\mathrm{trend}}=0.88$ vs. $0.90$ for unconstrained guidance) while reducing median off-target drift from $0.46$ to $0.20$ and 90th-percentile drift from $0.98$ to $0.33$. Drift magnitude tracks empirical target--off-target association, supporting the view that semantic instability is structured rather than incidental. A blinded human validation probe with radiology trainees further shows stronger agreement with intended progression orderings ($τ=0.61$ vs.\ $0.29$ for Pix2Pix). These results argue that medical image editing should be evaluated as trajectory-level semantic control, not as endpoint score maximization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。