arXiv:2608.27885cs.LG2026-08

双向扩散桥实现文本与图像互转,生成更灵活可靠。

There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation

论文配图:There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
图 1 · 摘自论文原文
  • 从文本直接插值到图像,路径显式包含源信息
  • 支持图像到文本逆向生成,统一双向流程
  • 适用于需要双向生成的多模态任务

多模态翻译(如文本到图像)是生成式AI的核心任务。现有方法存在两大局限:(1) 生成路径不直接反映源模态,限制采样算法灵活性;(2) 单向设计,无法实现反向生成(如图到文)。本文提出BIT:双向图像-文本扩散桥。与以往方法不同,BIT从文本出发,直接插值生成图像,提供(1) 源感知的生成路径,支持多样且灵活的采样算法;(2) 以终点为条件的可逆过程,可反向从图像到文本,构建统一的双向生成框架。BIT基于随机微分方程推导,具备可模拟的SDE形式和可扩展至高维的可计算损失函数。实验表明,BIT在性能上可比肩去噪扩散与确定性流基线,在多个视觉-语言及自然科学评估中表现更优。

原文摘要 · Abstract (English)

Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) follow generative paths that do not directly represent the source modality, limiting the flexibility of some sampling algorithms; and (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a source-aware generative path that enables diverse and flexible sampling algorithms; and (2) an endpoint-conditioned process that can be traversed from image to text, providing a unified, bidirectional generative framework. BIT is derived through stochastic calculus, yielding SDE forms amenable to simulation and tractable loss functions that scale to high dimensions. Our experiments show that BIT is competitive with denoising-diffusion and deterministic-flow baselines, and outperforms them on several vision--language and natural-science evaluations.

多模态生成扩散模型双向生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。