用文字生成3D沉浸式虚拟现实视频,无需实拍。
TexAVi: Generating Stereoscopic VR Video Clips from Text Descriptions
- 用文本生成图像,再用Stable Diffusion增强画面质量。
- 通过深度估计生成左右眼视图,拼接成立体视频。
- 适合虚拟现实内容快速生成,尤其适合影视与模拟场景。
尽管文生图、大语言模型和文生视频等生成模型已取得显著进展,但文生虚拟现实仍因训练数据不足及虚拟环境中实现真实深度与运动的复杂性而研究较少。本文提出一种整合现有生成系统的方法,从文本描述生成立体虚拟现实视频片段。整个流程分三步:首先使用基础文生图模型捕捉文本上下文;随后在生成的原始图像上应用Stable Diffusion,提升帧的逼真度与整体质量;最后利用深度估计算法生成左眼与右眼视图,拼接为并列画面以实现沉浸式观看体验。该方法对虚拟现实制作极具价值,可大幅减少实景拍摄与后期制作的时间投入。我们采用弗雷切特起始距离(Fréchet Inception Distance)与CLIP Score等图像评估技术,定量验证生成帧的质量,结果表明该方法具有优异性能。本工作展示了自然语言驱动图形在虚拟现实模拟等领域的巨大潜力。
原文摘要 · Abstract (English)
While generative models such as text-to-image, large language models and text-to-video have seen significant progress, the extension to text-to-virtual-reality remains largely unexplored, due to a deficit in training data and the complexity of achieving realistic depth and motion in virtual environments. This paper proposes an approach to coalesce existing generative systems to form a stereoscopic virtual reality video from text. Carried out in three main stages, we start with a base text-to-image model that captures context from an input text. We then employ Stable Diffusion on the rudimentary image produced, to generate frames with enhanced realism and overall quality. These frames are processed with depth estimation algorithms to create left-eye and right-eye views, which are stitched side-by-side to create an immersive viewing experience. Such systems would be highly beneficial in virtual reality production, since filming and scene building often require extensive hours of work and post-production effort. We utilize image evaluation techniques, specifically Fréchet Inception Distance and CLIP Score, to assess the visual quality of frames produced for the video. These quantitative measures establish the proficiency of the proposed method. Our work highlights the exciting possibilities of using natural language-driven graphics in fields like virtual reality simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。