让文生图模型精准控制镜头角度与焦距,提升画面表现力。
PreciseCam: Precise Camera Control for Text-to-Image Generation
- 仅用四个相机参数实现拍摄角度与镜头效果的精确控制。
- 在57,000张图像上验证了控制精度,优于传统提示工程。
- 适合需要影视级构图的艺术生成与设计场景。
图像作为艺术媒介常依赖特定镜头角度和镜头畸变来传达想法或情感,但现有文生图模型缺乏这种精细控制。本文提出一种高效通用的方法,可在生成摄影及艺术图像时实现精准相机控制。不同于依赖预设镜头类型的方法,本方法仅使用四个简单的相机外参与内参,无需预先存在的几何结构、参考3D物体或多视角数据。我们还构建了一个包含超过57,000张图像的新数据集,附带文本提示和真实相机参数。评估表明,该方法在文生图中实现了精确的相机控制,优于传统提示工程。相关数据、模型与代码已公开:https://graphics.unizar.es/projects/PreciseCam2024。
原文摘要 · Abstract (English)
Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that allows precise control over the camera when generating both photographic and artistic images. Unlike prior methods that rely on predefined shots, we rely solely on four simple extrinsic and intrinsic camera parameters, removing the need for pre-existing geometry, reference 3D objects, and multi-view data. We also present a novel dataset with more than 57,000 images, along with their text prompts and ground-truth camera parameters. Our evaluation shows precise camera control in text-to-image generation, surpassing traditional prompt engineering approaches. Our data, model, and code are publicly available at https://graphics.unizar.es/projects/PreciseCam2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。