让AI像专业摄影师一样精准操控镜头,懂物体体积和动态。
LensCraft: Your Professional Virtual Cinematographer
- 用仿真数据训练神经模型,兼顾创作意图与物体真实体积。
- 在动静场景中均实现高精度、高连贯的镜头运动,优于现有模型。
- 支持文本、轨迹、关键点等多方式控制,适合影视创作者使用。
数字创作者(从独立制片人到动画工作室)面临的核心挑战是将创意愿景转化为精确的镜头运动。尽管计算机视觉与人工智能取得进展,现有自动化拍摄系统仍受限于机械执行与创作意图之间的权衡。此前工作几乎都将主体简化为单一点,忽略其朝向与真实体积,严重削弱了拍摄中的空间感知能力。LensCraft通过数据驱动方法模拟专业摄影师的技巧,结合电影拍摄原则与实时动态场景适应能力,解决该问题。本方案采用专用仿真框架生成高质量训练数据,并构建先进神经模型,在忠实于剧本的同时,能感知主体体积与动态行为。系统支持多种输入模态(如文本提示、主体轨迹与体积、关键点或完整镜头路径),为创作者提供灵活的镜头引导工具。基于轻量级实时架构,LensCraft显著降低计算复杂度并提升推理速度,同时保持高质量输出。在静态与动态场景中的广泛评估显示,其准确性和连贯性达到前所未有的水平,树立智能摄像系统新基准。扩展结果、完整数据集、仿真环境、训练权重及源代码已公开于LensCraft网页。
原文摘要 · Abstract (English)
Digital creators, from indie filmmakers to animation studios, face a persistent bottleneck: translating their creative vision into precise camera movements. Despite significant progress in computer vision and artificial intelligence, current automated filming systems struggle with a fundamental trade-off between mechanical execution and creative intent. Crucially, almost all previous works simplify the subject to a single point-ignoring its orientation and true volume-severely limiting spatial awareness during filming. LensCraft solves this problem by mimicking the expertise of a professional cinematographer, using a data-driven approach that combines cinematographic principles with the flexibility to adapt to dynamic scenes in real time. Our solution combines a specialized simulation framework for generating high-fidelity training data with an advanced neural model that is faithful to the script while being aware of the volume and dynamic behavior of the subject. Additionally, our approach allows for flexible control via various input modalities, including text prompts, subject trajectory and volume, key points, or a full camera trajectory, offering creators a versatile tool to guide camera movements in line with their vision. Leveraging a lightweight real time architecture, LensCraft achieves markedly lower computational complexity and faster inference while maintaining high output quality. Extensive evaluation across static and dynamic scenarios reveals unprecedented accuracy and coherence, setting a new benchmark for intelligent camera systems compared to state-of-the-art models. Extended results, the complete dataset, simulation environment, trained model weights, and source code are publicly accessible on LensCraft Webpage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。