arXiv:2412.09323cs.CV2024-12被引 4

用文字提示生成立体视频,让3D效果制作更简单。

T-SVG: Text-Driven Stereoscopic Video Generation

  • 输入文字即可生成参考视频,转为双视角点云实现立体效果。
  • 无需训练,直接使用现有模型,支持快速更新。
  • 适合想轻松制作沉浸式视频的创作者和开发者。

立体视频的兴起为扩展现实(XR)和虚拟现实(VR)应用开辟了新前景,其沉浸式内容在多个平台吸引用户。然而,由于立体视差生成的技术复杂性,立体视频制作仍具挑战。视差指从两个不同视角观察物体时的位置差异,是实现深度感知的关键。为此,本文提出文本驱动立体视频生成(T-SVG)系统。该模型无关、零样本的方法通过文本提示生成参考视频,将其转化为3D点云序列,并从两个略有偏移的视角渲染,实现自然立体效果。T-SVG融合了先进的无训练文本到视频生成、深度估计与视频修复技术,架构灵活高效,可无缝集成新模型而无需重训。该系统简化了制作流程,使更多人能便捷创作立体视频,具有变革行业潜力。

原文摘要 · Abstract (English)

The advent of stereoscopic videos has opened new horizons in multimedia, particularly in extended reality (XR) and virtual reality (VR) applications, where immersive content captivates audiences across various platforms. Despite its growing popularity, producing stereoscopic videos remains challenging due to the technical complexities involved in generating stereo parallax. This refers to the positional differences of objects viewed from two distinct perspectives and is crucial for creating depth perception. This complex process poses significant challenges for creators aiming to deliver convincing and engaging presentations. To address these challenges, this paper introduces the Text-driven Stereoscopic Video Generation (T-SVG) system. This innovative, model-agnostic, zero-shot approach streamlines video generation by using text prompts to create reference videos. These videos are transformed into 3D point cloud sequences, which are rendered from two perspectives with subtle parallax differences, achieving a natural stereoscopic effect. T-SVG represents a significant advancement in stereoscopic content creation by integrating state-of-the-art, training-free techniques in text-to-video generation, depth estimation, and video inpainting. Its flexible architecture ensures high efficiency and user-friendliness, allowing seamless updates with newer models without retraining. By simplifying the production pipeline, T-SVG makes stereoscopic video generation accessible to a broader audience, demonstrating its potential to revolutionize the field.

立体视频文本生成3D渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。