用文字指令实时驱动3D高斯点云动画,免去建模和模拟步骤。
PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation
- 将3D动画转为文本驱动的4维场预测,直接操控高斯点云
- 单条文字提示即可实现颜色、透明度、位置的动态变化
- 适合初学者和专家,浏览器内即可实时生成特效
视觉特效是现代影视、游戏和AR/VR沉浸感的关键。传统3D特效制作需专业技能且耗时。现有生成方法多依赖计算量大的扩散模型,4D推理速度慢。本文将3D动画重构为场预测任务,提出一种文本驱动框架,通过大语言模型(LLMs)和视觉-语言模型(VLMs)生成时间变化的4D流场,作用于3D高斯点云。该方法可即时响应任意文本指令(如“让花瓶先变橙色,再爆炸”),实时更新高斯点云的颜色、透明度与位置。无需网格提取、手工或物理模拟,支持新手与专家在消费级设备甚至浏览器中快速创建体积场景动画。实验表明,简单文本指令即可生成引人注目的动态特效,显著降低绑定与高级建模的投入。本工作提供了一条快速、易用的语言驱动3D内容创作路径,有望进一步推动视觉特效的普及。代码已公开:https://obsphera.github.io/promptvfx/
原文摘要 · Abstract (English)
Visual effects (VFX) are key to immersion in modern films, games, and AR/VR. Creating 3D effects requires specialized expertise and training in 3D animation software and can be time consuming. Generative solutions typically rely on computationally intense methods such as diffusion models which can be slow at 4D inference. We reformulate 3D animation as a field prediction task and introduce a text-driven framework that infers a time-varying 4D flow field acting on 3D Gaussians. By leveraging large language models (LLMs) and vision-language models (VLMs) for function generation, our approach interprets arbitrary prompts (e.g., "make the vase glow orange, then explode") and instantly updates color, opacity, and positions of 3D Gaussians in real time. This design avoids overheads such as mesh extraction, manual or physics-based simulations and allows both novice and expert users to animate volumetric scenes with minimal effort on a consumer device even in a web browser. Experimental results show that simple textual instructions suffice to generate compelling time-varying VFX, reducing the manual effort typically required for rigging or advanced modeling. We thus present a fast and accessible pathway to language-driven 3D content creation that can pave the way to democratize VFX further. Code available at https://obsphera.github.io/promptvfx/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。