arXiv:2503.08434cs.GRcs.CV2025-03SIGGRAPH被引 11

让文生图模型像相机一样精准控制虚化程度,保持画面内容不变。

Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

  • 通过物理模糊参数显式控制扩散模型的虚化强度。
  • 在真实图像与合成模糊数据上联合训练,实现内容与模糊分离。
  • 支持图像编辑和跨架构通用,虚化调节更自然可控。

近期大规模文生图模型虽显著推动创意领域发展,但当前扩散模型依赖提示词工程模拟景深效果,常导致视觉失真且改变场景内容。本文提出Bokeh Diffusion,一种场景一致的虚化控制框架,通过物理定义的散焦模糊参数显式条件化扩散模型。针对真实多镜头图像对稀缺问题,引入混合训练流程,将真实图像与合成模糊增强对齐,提供多样场景与监督信号以学习内容与镜头模糊的解耦。核心是基于同一场景不同虚化程度图像对训练的接地自注意力机制,可双向调整模糊强度并保持原始场景不变。大量实验表明,该方法实现灵活、类镜头的虚化控制,支持逆向图像编辑等下游应用,并在Stable Diffusion与FLUX架构上具有良好泛化性。

原文摘要 · Abstract (English)

Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textual prompts; however, while traditional photography offers precise control over camera settings to shape visual aesthetics - such as depth-of-field via aperture - current diffusion models typically rely on prompt engineering to mimic such effects. This approach often results in crude approximations and inadvertently alters the scene content. In this work, we propose Bokeh Diffusion, a scene-consistent bokeh control framework that explicitly conditions a diffusion model on a physical defocus blur parameter. To overcome the scarcity of paired real-world images captured under different camera settings, we introduce a hybrid training pipeline that aligns in-the-wild images with synthetic blur augmentations, providing diverse scenes and subjects as well as supervision to learn the separation of image content from lens blur. Central to our framework is our grounded self-attention mechanism, trained on image pairs with different bokeh levels of the same scene, which enables blur strength to be adjusted in both directions while preserving the underlying scene. Extensive experiments demonstrate that our approach enables flexible, lens-like blur control, supports downstream applications such as real image editing via inversion, and generalizes effectively across both Stable Diffusion and FLUX architectures.

文生图虚化控制扩散模型图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。