arXiv:2411.02394cs.CV2024-11被引 32

用自然语言指令自动生成逼真动态视频特效

AutoVFX: Physically Realistic Video Editing from Natural Language Instructions

  • 融合神经场景建模与物理仿真,实现语言驱动的视觉效果生成
  • 在生成质量、指令对齐和物理合理性上显著优于现有方法
  • 适合影视创作、游戏开发等需要快速生成特效的场景

现代视觉特效软件虽能生成几乎任意图像,但制作过程仍繁琐复杂,普通用户难以使用。本文提出AutoVFX框架,仅需一段视频和自然语言指令,即可自动生成具有物理真实感的动态视觉特效。通过整合神经场景建模、基于大语言模型的代码生成与物理模拟,AutoVFX实现了以自然语言直接控制的逼真编辑效果。我们在多种视频与指令组合上进行了广泛实验,定量与定性结果表明,AutoVFX在生成质量、指令对齐性、编辑灵活性及物理合理性方面均大幅超越现有方法。

原文摘要 · Abstract (English)

Modern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility.

视频生成语言控制物理仿真AI特效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。