用文本指令自动生成无人机航拍轨迹,实现自然语言驱动的电影级航拍。
DiffusionCinema: Text-to-Aerial Cinematography
- 输入自然语言描述,系统结合视觉快照生成符合场景与语义的飞行路径。
- 用户实验显示工作负荷降低63%,心理压力和挫败感显著下降。
- 适合影视创作、摄影爱好者,无需操控技能即可实现专业级航拍。
我们提出一种新型无人机辅助创意拍摄系统,利用扩散模型解析高层自然语言指令,并自动生成最优飞行轨迹以完成电影级视频录制。用户无需手动操控无人机,仅需描述期望镜头(如“从右侧缓慢环绕我并揭示背景瀑布”)。系统将提示词与机载相机的初始视觉快照编码,通过扩散模型采样满足场景几何结构与镜头语义的合理时空运动规划。生成的飞行轨迹由无人机自主执行,记录出平滑且可重复的视频片段。用户评估使用NASA-TLX量表显示,本系统整体工作负荷均值为21.6,远低于传统遥控器的58.1;心理负荷(11.5 vs. 60.5)与挫败感(14.0 vs. 54.5)也显著更低,验证了在自主文本驱动飞行控制中的明显可用性优势。本研究展示了一种新交互范式:文本到电影航拍,其中扩散模型充当‘创意操作员’,直接将故事意图转化为空中运动。
原文摘要 · Abstract (English)
We propose a novel Unmanned Aerial Vehicles (UAV) assisted creative capture system that leverages diffusion models to interpret high-level natural language prompts and automatically generate optimal flight trajectories for cinematic video recording. Instead of manually piloting the drone, the user simply describes the desired shot (e.g., "orbit around me slowly from the right and reveal the background waterfall"). Our system encodes the prompt along with an initial visual snapshot from the onboard camera, and a diffusion model samples plausible spatio-temporal motion plans that satisfy both the scene geometry and shot semantics. The generated flight trajectory is then executed autonomously by the UAV to record smooth, repeatable video clips that match the prompt. User evaluation using NASA-TLX showed a significantly lower overall workload with our interface (M = 21.6) compared to a traditional remote controller (M = 58.1), demonstrating a substantial reduction in perceived effort. Mental demand (M = 11.5 vs. 60.5) and frustration (M = 14.0 vs. 54.5) were also markedly lower for our system, confirming clear usability advantages in autonomous text-driven flight control. This project demonstrates a new interaction paradigm: text-to-cinema flight, where diffusion models act as the "creative operator" converting story intentions directly into aerial motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。