用文本生成立体图像,靠微调Stable Diffusion实现
Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards
- 基于Stable Diffusion微调,利用其预训练先验知识
- 引入一致性奖励机制,提升立体图像与文本对齐度
- 适合需要高质量立体图像生成的研究者和开发者
本文提出一种基于扩散模型的新方法,仅凭文本提示即可生成立体图像。由于大视差立体图像数据集稀缺,直接从零训练扩散模型不可行。为此,我们借助Stable Diffusion已学习的强先验知识,在立体图像数据集上进行微调,使其适应立体生成任务。为进一步提升立体一致性与文本-图像对齐效果,我们采用提示对齐策略,并引入自研的立体一致性奖励函数进行优化。大量实验表明,该方法在多种场景下均能生成高质量立体图像,优于现有方法。
原文摘要 · Abstract (English)
In this paper, we propose a novel diffusion-based approach to generate stereo images given a text prompt. Since stereo image datasets with large baselines are scarce, training a diffusion model from scratch is not feasible. Therefore, we propose leveraging the strong priors learned by Stable Diffusion and fine-tuning it on stereo image datasets to adapt it to the task of stereo generation. To improve stereo consistency and text-to-image alignment, we further tune the model using prompt alignment and our proposed stereo consistency reward functions. Comprehensive experiments demonstrate the superiority of our approach in generating high-quality stereo images across diverse scenarios, outperforming existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。