arXiv:2511.00801cs.CVcs.MM2025-11被引 2

通过成功与失败轨迹训练医疗图像编辑,提升生成结果的临床合理性。

Med-Banana: Learning Quality-Controlled Medical Image Editing from Success-and-Failure Trajectories

  • 利用编辑过程中的成功与失败轨迹进行联合训练
  • 在多个评估中显著优于现有医学图像编辑器
  • 适合需要高临床可信度的医疗AI研究者

文本引导的医学图像编辑需满足病灶要求,同时保持解剖结构、模态特异性外观和临床合理性。然而,现有数据集仅以最终接受的编辑为监督信号,忽略生成过程中产生的失败尝试。我们提出,这些失败案例提供了关键的质量控制信息:说明什么应被拒绝、为何编辑不合法或视觉不合理,以及如何优化指令。为此,我们构建了Med-Banana-80K,一个包含80,000条成功与失败编辑轨迹的大规模资源,涵盖候选图像、验证结果、拒绝原因及提示修正。基于此,Med-Banana联合训练编辑器、验证器与修正器,实现从接受与拒绝尝试中学习的编辑-验证-修正推理流程。在多类评测中,包括多模态大模型判别、盲评专家评估、源图保留性与真实-合成分离性探测,均表现出一致优越性。代码与数据已公开。

原文摘要 · Abstract (English)

Text-guided medical image editing must satisfy the requested pathology while preserving anatomy, modality-specific appearance, and clinical plausibility. However, existing datasets largely supervise editors with final accepted edits and discard the failed attempts produced during generation. We argue that these failures provide essential supervision for quality control: they specify what should be rejected, why an edit is medically or visually invalid, and how the instruction should be revised. We present Med-Banana, a trajectory-supervised framework for quality-controlled medical image editing. We introduce Med-Banana-80K, a large-scale resource of success-and-failure editing trajectories with candidate images, verification outcomes, rejection reasons, and prompt refinements. Building on it, Med-Banana jointly trains an editor, verifier, and refiner, enabling edit--verify--refine inference from accepted and rejected attempts. Experiments across MLLM judges, blind expert assessment, source-preservation and real--synthetic separability probes demonstrate consistent improvements over open medical image editors. Code and data are publicly available.

医疗图像编辑质量控制轨迹学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。