arXiv:2601.08095cs.CV2026-01

用扩散模型自动生成适配真实场景的合成数据集

From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models

  • 通过可控修复生成特定背景下的目标物体
  • 多模态评估确保生成数据兼具准确与美观
  • 融合用户偏好筛选,适合实际部署需求

本文提出一种基于扩散模型的自动化流程,用于生成领域特定的合成数据集,以缓解预训练模型与真实部署环境间的分布偏移问题。该三阶段框架首先在特定领域背景中通过可控修复合成目标物体;生成结果经多模态评估,包括目标检测性能、美学评分及视觉-语言对齐度;最后利用用户偏好分类器捕捉主观选择标准。该流程可高效构建高质量、可部署的数据集,显著减少对大量真实数据采集的依赖。

原文摘要 · Abstract (English)

In this paper, we present an automated pipeline for generating domain-specific synthetic datasets with diffusion models, addressing the distribution shift between pre-trained models and real-world deployment environments. Our three-stage framework first synthesizes target objects within domain-specific backgrounds through controlled inpainting. The generated outputs are then validated via a multi-modal assessment that integrates object detection, aesthetic scoring, and vision-language alignment. Finally, a user-preference classifier is employed to capture subjective selection criteria. This pipeline enables the efficient construction of high-quality, deployable datasets while reducing reliance on extensive real-world data collection.

扩散模型数据合成领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。