arXiv:2509.18796cs.CV2025-09

让生成的手术图像更符合下游任务需求,提升模型性能。

Towards Application Aligned Synthetic Surgical Image Synthesis

  • 用偏好对比对训练扩散模型,使其生成更适配下游任务的图像。
  • 在3个数据集上分类任务提升7%~9%,分割任务提升2%~10%。
  • 特别改善小样本类表现,适合医疗视觉与生成模型结合研究者。

标注手术数据的匮乏严重制约了计算机辅助干预中深度学习系统的发展。尽管扩散模型能生成逼真图像,但常因数据记忆导致样本不一致或缺乏多样性,可能损害下游性能。我们提出 extit{Surgical Application-Aligned Diffusion}(SAADi)框架,将扩散模型生成过程与下游模型偏好对齐。通过构建偏好与非偏好合成图像对,并对扩散模型进行轻量微调,显式对齐生成目标。在三个手术数据集上的实验显示,分类任务性能提升7%~9%,分割任务提升2%~10%,尤其在低频类别上改善显著。迭代优化合成样本可进一步提升4%~10%。相比基线方法,本方法克服样本退化问题,确立任务感知对齐为缓解数据稀缺的关键原则,推动手术视觉应用发展。

原文摘要 · Abstract (English)

The scarcity of annotated surgical data poses a significant challenge for developing deep learning systems in computer-assisted interventions. While diffusion models can synthesize realistic images, they often suffer from data memorization, resulting in inconsistent or non-diverse samples that may fail to improve, or even harm, downstream performance. We introduce \emph{Surgical Application-Aligned Diffusion} (SAADi), a new framework that aligns diffusion models with samples preferred by downstream models. Our method constructs pairs of \emph{preferred} and \emph{non-preferred} synthetic images and employs lightweight fine-tuning of diffusion models to align the image generation process with downstream objectives explicitly. Experiments on three surgical datasets demonstrate consistent gains of $7$--$9\%$ in classification and $2$--$10\%$ in segmentation tasks, with the considerable improvements observed for underrepresented classes. Iterative refinement of synthetic samples further boosts performance by $4$--$10\%$. Unlike baseline approaches, our method overcomes sample degradation and establishes task-aware alignment as a key principle for mitigating data scarcity and advancing surgical vision applications.

手术图像扩散模型数据生成医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。