arXiv:2509.16176cs.RO2025-09被引 2

让无人机通过对话自动完成高质量航拍,无需专业操作。

Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories

  • 用大模型理解自然语言指令,自动生成拍摄路径。
  • 结合美学反馈优化视角,生成符合专业标准的镜头。
  • 适合影视创作、智能摄影等无技术门槛场景使用。

我们提出一种名为ACDC(Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories)的自主无人机航拍系统,通过人机自然语言对话实现无人机航拍。传统方法依赖人工设定航点与视角,耗时且效果不一。本文利用大语言模型(LLMs)和视觉基础模型(VFMs),将自由形式的自然语言提示直接转化为可执行的室内无人机视频巡演。方法包括:基于视觉-语言检索的初始航点选择、基于偏好反馈的贝叶斯优化框架用于姿态精修,以及生成安全四旋翼轨迹的运动规划器。通过仿真与软硬件联合实验验证,ACDC在多种室内场景下均能鲁棒生成专业级画面,无需用户具备机器人或影视制作知识。结果表明,具身AI代理有望实现从开放词汇对话到真实世界自主航拍的闭环。

原文摘要 · Abstract (English)

We present Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories (ACDC), an autonomous drone cinematography system driven by natural language communication between human directors and drones. The main limitation of previous drone cinematography workflows is that they require manual selection of waypoints and view angles based on predefined human intent, which is labor-intensive and yields inconsistent performance. In this paper, we propose employing large language models (LLMs) and vision foundation models (VFMs) to convert free-form natural language prompts directly into executable indoor UAV video tours. Specifically, our method comprises a vision-language retrieval pipeline for initial waypoint selection, a preference-based Bayesian optimization framework that refines poses using aesthetic feedback, and a motion planner that generates safe quadrotor trajectories. We validate ACDC through both simulation and hardware-in-the-loop experiments, demonstrating that it robustly produces professional-quality footage across diverse indoor scenes without requiring expertise in robotics or cinematography. These results highlight the potential of embodied AI agents to close the loop from open-vocabulary dialogue to real-world autonomous aerial cinematography.

无人机航拍自然语言控制具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。