arXiv:2506.05367cs.CV2025-06被引 2

用文本生成立体图像,靠微调Stable Diffusion实现

Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards

  • 基于Stable Diffusion微调,利用其预训练先验知识
  • 引入一致性奖励机制,提升立体图像与文本对齐度
  • 适合需要高质量立体图像生成的研究者和开发者

本文提出一种基于扩散模型的新方法,仅凭文本提示即可生成立体图像。由于大视差立体图像数据集稀缺,直接从零训练扩散模型不可行。为此,我们借助Stable Diffusion已学习的强先验知识,在立体图像数据集上进行微调,使其适应立体生成任务。为进一步提升立体一致性与文本-图像对齐效果,我们采用提示对齐策略,并引入自研的立体一致性奖励函数进行优化。大量实验表明,该方法在多种场景下均能生成高质量立体图像,优于现有方法。

原文摘要 · Abstract (English)

In this paper, we propose a novel diffusion-based approach to generate stereo images given a text prompt. Since stereo image datasets with large baselines are scarce, training a diffusion model from scratch is not feasible. Therefore, we propose leveraging the strong priors learned by Stable Diffusion and fine-tuning it on stereo image datasets to adapt it to the task of stereo generation. To improve stereo consistency and text-to-image alignment, we further tune the model using prompt alignment and our proposed stereo consistency reward functions. Comprehensive experiments demonstrate the superiority of our approach in generating high-quality stereo images across diverse scenarios, outperforming existing methods.

立体生成扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。