arXiv:2506.17450cs.CVcs.GR2025-06被引 9

用3D控制实现图像合成,可灵活替换背景或调整物体与镜头关系。

BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing

  • 分层-编辑-融合三步流程,用3D实体操控图像元素。
  • 在复杂场景合成任务中效果优于现有方法,支持背景替换和动态调整。
  • 适合需要精细视觉合成的设计师、影视制作人员使用。

我们提出BlenderFusion,一种生成式视觉合成框架,通过重组物体、相机和背景来生成新场景。其采用分层-编辑-合成流水线:(i) 将视觉输入分割并转换为可编辑的3D实体(分层);(ii) 在Blender中进行3D对齐控制的编辑(编辑);(iii) 使用生成式合成器将原始场景与编辑后场景并行融合(合成)。生成式合成器基于预训练扩散模型,通过两种关键训练策略微调:(i) 源图掩码,支持灵活修改如背景替换;(ii) 模拟物体抖动,实现物体与相机的解耦控制。BlenderFusion在复杂组合场景编辑任务中显著优于先前方法。

原文摘要 · Abstract (English)

We present BlenderFusion, a generative visual compositing framework that synthesizes new scenes by recomposing objects, camera, and background. It follows a layering-editing-compositing pipeline: (i) segmenting and converting visual inputs into editable 3D entities (layering), (ii) editing them in Blender with 3D-grounded control (editing), and (iii) fusing them into a coherent scene using a generative compositor (compositing). Our generative compositor extends a pre-trained diffusion model to process both the original (source) and edited (target) scenes in parallel. It is fine-tuned on video frames with two key training strategies: (i) source masking, enabling flexible modifications like background replacement; (ii) simulated object jittering, facilitating disentangled control over objects and camera. BlenderFusion significantly outperforms prior methods in complex compositional scene editing tasks.

视觉合成3D控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。