让摄像机轨迹生成更懂导演审美,自动优化构图与视觉效果。
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation

- 用实时渲染+视觉语言模型评估生成镜头的视觉偏好
- 将角色出框率从38%降到近乎0%,保持运动几何精度
- 适合影视创作、动画生成中需要高质量镜头设计的场景
电影级摄像机控制依赖导演与摄影师之间的紧密反馈循环,持续调整镜头运动与构图。现有生成式摄像机系统虽能生成多样化的文本驱动轨迹,但缺乏“导演在环”的反馈机制,且未显式优化画面是否美观。导致生成结果虽在运动分布内,却存在构图不佳、角色出画、视觉美感差等问题。本文提出VERTIGO,首个针对摄像机轨迹生成器的视觉偏好优化框架。该框架利用Unity实时图形引擎渲染2D预览画面,通过一个经过电影美学微调的视觉-语言模型,结合提出的循环语义相似性机制,对渲染画面与文本提示进行对齐评分。该评分作为直接偏好优化(DPO)后训练的视觉偏好信号。定量评估与用户研究(基于Unity渲染及扩散模型相机到视频流水线)均显示,该方法在条件遵循度、构图质量与感知真实感上显著提升。特别地,角色出框率从38%降至接近0%,同时保持摄像机运动的几何保真度。用户偏好研究进一步证实,参与者在构图、一致性、提示遵循与美学质量方面更青睐VERTIGO方案。
原文摘要 · Abstract (English)
Cinematic camera control relies on a tight feedback loop between director and cinematographer, where camera motion and framing are continuously reviewed and refined. Recent generative camera systems can produce diverse, text-conditioned trajectories, but they lack this "director in the loop" and have no explicit supervision of whether a shot is visually desirable. This results in in-distribution camera motion but poor framing, off-screen characters, and undesirable visual aesthetics. In this paper, we introduce VERTIGO, the first framework for visual preference optimization of camera trajectory generators. Our framework leverages a real-time graphics engine (Unity) to render 2D visual previews from generated camera motion. A cinematically fine-tuned vision-language model then scores these previews using our proposed cyclic semantic similarity mechanism, which aligns renders with text prompts. This process provides the visual preference signals for Direct Preference Optimization (DPO) post-training. Both quantitative evaluations and user studies on Unity renders and diffusion-based Camera-to-Video pipelines show consistent gains in condition adherence, framing quality, and perceptual realism. Notably, VERTIGO reduces the character off-screen rate from 38% to nearly 0% while preserving the geometric fidelity of camera motion. User study participants further prefer VERTIGO over baselines across composition, consistency, prompt adherence, and aesthetic quality, confirming the perceptual benefits of our visual preference post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。