arXiv:2601.18623cs.CV2026-01中稿 · ICLR被引 4

让扩散模型按需调整跨模态转换,更准更快。

Adaptive Domain Shift in Diffusion Models for Cross-Modality Image Translation

论文配图:Adaptive Domain Shift in Diffusion Models for Cross-Modality Image Translation
图 1 · 摘自论文原文
  • 在反向采样每一步动态预测空间混合场,实现局部修正
  • 减少无效路径搜索,结构保真度提升,收敛步数减少
  • 适合医学影像、遥感等需要精准跨模态转换的场景

跨模态图像转换仍存在脆弱且低效的问题。标准扩散方法通常依赖单一全局线性域转移,我们发现这种简化会迫使采样器进入离流形的高代价区域,增加校正负担并引发语义漂移。我们称此共性缺陷为固定时序域转移。本文将域转移动态直接嵌入生成过程:在每个反向步骤预测空间变化的混合场,并注入目标一致的显式恢复项以引导漂移。这种步骤内指导使大幅更新保持在流形上,将模型角色从全局对齐转为局部残差修正。我们提出连续时间形式及精确解,推导出保持边际一致性的实用一阶采样器。实验表明,在医学影像、遥感与电致发光语义映射任务中,本框架显著提升结构保真度与语义一致性,同时减少去噪步数。

原文摘要 · Abstract (English)

Cross-modal image translation remains brittle and inefficient. Standard diffusion approaches often rely on a single, global linear transfer between domains. We find that this shortcut forces the sampler to traverse off-manifold, high-cost regions, inflating the correction burden and inviting semantic drift. We refer to this shared failure mode as fixed-schedule domain transfer. In this paper, we embed domain-shift dynamics directly into the generative process. Our model predicts a spatially varying mixing field at every reverse step and injects an explicit, target-consistent restoration term into the drift. This in-step guidance keeps large updates on-manifold and shifts the model's role from global alignment to local residual correction. We provide a continuous-time formulation with an exact solution form and derive a practical first-order sampler that preserves marginal consistency. Empirically, across translation tasks in medical imaging, remote sensing, and electroluminescence semantic mapping, our framework improves structural fidelity and semantic consistency while converging in fewer denoising steps.

扩散模型跨模态图像转换医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。