arXiv:2504.09472cs.CV2025-04

仅用一张图片和一段视频,就能零样本生成有电影感的镜头运动。

Zero-Shot Personalized Camera Motion Control for Image-to-Video Synthesis

  • 通过双低秩适配网络优化,分离图像外观与动态运动。
  • 用户研究显示90.45%偏好其运镜准确性,70.31%认可场景保持度。
  • 无需3D数据或复杂界面,适合影视新手和教育创作者使用。

非专业创作者在使用生成工具时,难以精准描述细腻且富有表现力的镜头运动,形成“表达鸿沟”,使通用文本提示无法实现电影化构想。为解决这一问题,本文提出一种基于扩散模型的零样本个性化镜头运动控制框架,仅需单个参考视频即可将电影级运镜迁移至用户提供的静态图像,无需3D数据、预设轨迹或复杂图形界面。核心技术包括推理阶段优化策略,采用双低秩适配(LoRA)网络并引入正交性正则项,促进空间外观与时间运动更新的解耦;结合基于单应性的精修策略,提供弱几何引导。我们引入新评估指标CameraScore,并开展两项用户研究:72人感知实验表明,本方法在运镜准确性(90.45%偏好)和场景保留度(70.31%偏好)上显著优于基线;12人任务型交互研究证实,相比标准文本或预设提示,本工作在可用性与创作控制力方面均有显著提升(p < 0.001)。本研究为跨场景镜头运动迁移奠定基础。

原文摘要 · Abstract (English)

Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts fail to capture cinematic vision. This barrier limits individual creativity and restricts the accessibility of cinematic production for small-scale industries and educational content creators. To address this, we present a zero-shot diffusion-based framework for personalized camera motion control, enabling the transfer of cinematic movements from a single reference video onto a user-provided static image without requiring 3D data, predefined trajectories, or complex graphical interfaces. Our technical contribution involves an inference-time optimization strategy using dual Low-Rank Adaptation (LoRA) networks, with an orthogonality regularizer that encourages separation between spatial appearance and temporal motion updates, alongside a homography-based refinement strategy that provides weak geometric guidance. We evaluate our approach using a new metric, CameraScore, and two distinct user studies. A 72-participant perceptual study demonstrates that our method significantly outperforms existing baselines in motion accuracy (90.45% preference) and scene preservation (70.31% preference). Furthermore, a 12-participant task-based interaction study confirms that our workflow significantly improves usability and creative control (p < 0.001) compared to standard text- or preset-based prompts. We hope this work lays a foundation for future advancements in camera motion transfer across diverse scenes.

视频生成扩散模型镜头控制零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。