让全景视频生成可拖拽控制,实现精准场景与物体运动操控
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
- 通过球面运动估计器和控制模块,支持拖拽式全景视频生成
- 在新数据集Move360上实现高质量全景视频生成,运动更自然
- 适合虚拟现实、沉浸式内容创作人群,提升动态场景控制精度
随着虚拟现实日益普及,对可控制的沉浸式动态全景视频(ODVs)生成需求不断增长。现有文本到全景视频生成方法虽效果出色,但仅依赖文本输入易导致内容不准确和不一致。尽管近期运动控制技术可实现细粒度视频生成控制,但直接应用于全景视频常引发空间畸变且性能不佳,尤其在复杂球面运动下。为此,我们提出OmniDrag,首个实现场景级与物体级运动控制的全景图像到视频生成方法。基于预训练视频扩散模型,引入球面控制模块,并与时间注意力层联合微调以处理复杂球面运动。同时开发新型球面运动估计算法,可准确提取运动控制信号,支持用户通过绘制控制点实现拖拽式生成。此外,构建新数据集Move360,解决大范围场景与物体运动的全景视频数据稀缺问题。实验表明,OmniDrag在整体场景与细粒度物体控制方面显著优于现有方法。
原文摘要 · Abstract (English)
As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, they struggle with content inaccuracies and inconsistencies due to reliance solely on textual inputs. Although recent motion control techniques provide fine-grained control for video generation, directly applying these methods to ODVs often results in spatial distortion and unsatisfactory performance, especially with complex spherical motions. To tackle these challenges, we propose OmniDrag, the first approach enabling both scene- and object-level motion control for accurate, high-quality omnidirectional image-to-video generation. Building on pretrained video diffusion models, we introduce an omnidirectional control module, which is jointly fine-tuned with temporal attention layers to effectively handle complex spherical motion. In addition, we develop a novel spherical motion estimator that accurately extracts motion-control signals and allows users to perform drag-style ODV generation by simply drawing handle and target points. We also present a new dataset, named Move360, addressing the scarcity of ODV data with large scene and object motions. Experiments demonstrate the significant superiority of OmniDrag in achieving holistic scene-level and fine-grained object-level control for ODV generation. The project page is available at https://lwq20020127.github.io/OmniDrag.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。