让合成图像更真实:动态调节3D结构引导,避免过度约束
AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation

- 用自监督模块动态调整控制网络强度,防止过度引导
- 生成图像在2D和3D标注上都更准确,提升数据质量
- 适合需要高质量合成数据的视觉任务研究者
合成数据生成已成为提升计算机视觉数据可扩展性的有力工具。基于扩散模型的方法已展现出强大的逼真度。然而,在生成图像中保持精确的3D结构与姿态一致性仍是挑战。现有方法依赖边缘图等视觉提示引导扩散模型,但常出现过强引导导致的失真,降低图像真实感并限制数据集质量。本文提出一种基于扩散模型的图像生成框架——自适应3D感知合成数据生成(AC3S),通过引入自监督视觉提示调制器,动态调节ControlNet的条件强度,避免过强引导,使扩散模型保持生成表达力。为进一步提升多样性与语义一致性,我们构建多智能体视觉语言模型框架,生成与底层几何结构一致的详细、3D感知提示。两者结合实现了高保真合成数据集的可扩展生成,包含精确的2D与3D标注。大量实验表明,该方法显著提升图像质量与下游任务性能。
原文摘要 · Abstract (English)
Synthetic data generation has emerged as a powerful tool for improving data scalability in computer vision. Recent diffusion-based pipelines have demonstrated strong photorealism. However, how to enforce precise 3D structure and pose consistency in generated images remains challenging. Existing methods leverage visual prompts such as edge maps to guide diffusion models, but often suffer from over-conditioning artifacts that degrade image realism and limit dataset quality. In this paper, we present a diffusion-based image generation framework that enforces 3D structural alignment while preserving photorealism through adaptive conditioning. Our framework, Adaptive Conditioning for 3D-Aware Synthetic Data Generation (AC3S), introduces a self-supervised visual prompt modulator that dynamically adjusts the strength of ControlNet conditioning, preventing over-conditioning and enabling the diffusion model to retain its generative expressiveness. To further enhance diversity and semantic consistency, we develop a multi-agent vision language model framework that composes detailed and 3D-aware prompts aligned with the underlying geometric structure. Together, these components enable the scalable generation of high-quality synthetic datasets with accurate 2D and 3D annotations. Extensive experiments demonstrate that our method significantly improves image quality and downstream utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。