arXiv:2412.02168cs.CV2024-12CVPR被引 28

让AI生成照片时能真实模拟不同镜头视角,保持场景一致。

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

  • 通过提升维度与学习镜头差异,实现镜头参数的精准控制。
  • 在24mm与70mm镜头对比下,生成图像更符合真实物理透视。
  • 适合专业摄影与需要镜头控制的图像生成场景。

当前图像生成技术虽能从文本生成较逼真的图像,但当要求生成特定相机设置(如24mm与70mm镜头)下的不同视场角时,模型无法正确理解并生成场景一致的图像。这一局限不仅阻碍了生成式工具在专业摄影中的应用,也暴露了数据驱动模型与真实物理设定之间的脱节。本文提出Generative Photography框架,支持在内容生成过程中控制相机内参。核心创新在于维度提升与差分相机内参学习,实现了不同相机设置间的平滑、一致过渡。实验表明,该方法在生成场景一致性与逼真度上显著优于Stable Diffusion 3和FLUX等先进模型。代码与更多结果见https://generative-photography.github.io/project。

原文摘要 · Abstract (English)

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the generator will not be able to interpret and generate scene-consistent images. This limitation not only hinders the adoption of generative tools in professional photography but also highlights the broader challenge of aligning data-driven models with real-world physical settings. In this paper, we introduce Generative Photography, a framework that allows controlling camera intrinsic settings during content generation. The core innovation of this work are the concepts of Dimensionality Lifting and Differential Camera Intrinsics Learning, enabling smooth and consistent transitions across different camera settings. Experimental results show that our method produces significantly more scene-consistent photorealistic images than state-of-the-art models such as Stable Diffusion 3 and FLUX. Our code and additional results are available at https://generative-photography.github.io/project.

图像生成相机控制真实感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。