arXiv:2604.09022cs.CV2026-04

用3D路径追踪生成高质量图像对,解决扩散模型训练中的数据伪影问题。

BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training

论文配图:BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training
图 1 · 摘自论文原文
  • 基于3D场景路径追踪生成合成图像,结合物体中心相机布局
  • 构建的FineBLEND数据集在视觉质量上优于多个主流数据集
  • 开源框架支持自定义3D场景数据集生成,适合研究者扩展应用

随着扩散模型的快速发展,合成数据生成成为满足大规模图像数据需求的有前景方法。然而,纯扩散模型生成的图像常存在视觉不一致问题,基于此类数据训练会引发自噬反馈循环,导致模型崩溃,即模型自噬紊乱(MAD)。为此,我们提出BlendFusion,一种基于路径追踪的可扩展3D场景合成数据生成框架。该流程包含物体中心相机布置策略、鲁棒过滤机制和自动图文标注,用于生成高质量图像-文本对。利用该流程,我们构建了由多样化3D场景组成的FineBLEND图像-文本数据集。我们对FineBLEND的质量进行了实证分析,并与多个广泛使用数据集进行对比。同时,我们展示了物体中心相机布置策略相较于无物体感知采样方法的优势。我们的开源框架具备高度可配置性,使社区能够从3D场景中生成自有数据集。

原文摘要 · Abstract (English)

With the rapid adoption of diffusion models, synthetic data generation has emerged as a promising approach for addressing the growing demand for large-scale image datasets. However, images generated purely by diffusion models often exhibit visual inconsistencies, and training models on such data can create an autophagous feedback loop that leads to model collapse, commonly referred to as Model Autophagy Disorder (MAD). To address these challenges, we propose BlendFusion, a scalable framework for synthetic data generation from 3D scenes using path tracing. Our pipeline incorporates an object-centric camera placement strategy, robust filtering mechanisms, and automatic captioning to produce high-quality image-caption pairs. Using this pipeline, we curate FineBLEND, an image-caption dataset constructed from a diverse set of 3D scenes. We empirically analyze the quality of FineBLEND and compare it to several widely used image-caption datasets. We also demonstrate the effectiveness of our object-centric camera placement strategy relative to object-agnostic sampling approaches. Our open-source framework is designed for high configurability, enabling the community to create their own datasets from 3D scenes.

扩散模型合成数据3D生成图像-文本对

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。