arXiv:2605.09699eess.IVcs.CV2026-05

用可控生成与多阶段筛选,让合成数据有效提升低数据场景性能。

A Real-Calibrated Synthetic-First Data Engine

论文配图:A Real-Calibrated Synthetic-First Data Engine
图 1 · 摘自论文原文
  • 合成图像通过可控扩散生成,再经多阶段过滤与选择。
  • 结合真实数据锚点时,合成数据显著提升模型性能。
  • 适合需要低成本标注的视觉任务开发者使用。

现代计算机视觉系统在数据稀缺领域面临性能瓶颈,因大规模高质量标注数据收集成本高或不切实际。尽管可控扩散模型可实现可扩展的合成图像生成,但直接使用合成数据增强常因数据集级质量差和反馈机制不足导致性能不稳定。本文提出一种真实校准的合成优先数据引擎,是一个模块化数据工程框架,将可控扩散生成与多阶段筛选/过滤集成于统一流程中,支持不确定性驱动选择与人工验证。该方法不引入新生成算法,而是聚焦于系统性数据构建,以提升低数据场景下合成数据增强的实际可靠性。框架采用基于CLI的模块化设计,生成、过滤、选择与验证组件可独立配置替换,强调可复现性、灵活性及真实工作流部署。在人体姿态估计任务上的实证评估表明:当合成数据作为近零人工标注成本的增强数据与真实锚点结合时,能显著提升真实数据基线表现;而纯合成数据训练仍远低于纯真实数据训练性能。补充的分割诊断显示相同域差距模式。结果凸显了面向低数据增强的数据中心化编排的实用价值。

原文摘要 · Abstract (English)

Modern computer vision systems increasingly encounter performance limitations in data-scarce domains, where collecting large-scale, high-quality labeled data is costly or impractical. While controllable diffusion models enable scalable synthetic image generation, directly applying synthetic augmentation often leads to unstable performance gains due to dataset-level quality issues and insufficient feedback mechanisms. In this work, we present a Real-Calibrated Synthetic-First Data Engine, a modular data engineering framework that combines controllable diffusion generation and multi-stage curation/filtering within a unified pipeline, with optional support for uncertainty-driven selection and human verification. Instead of introducing new generative algorithms, our approach focuses on systematic dataset construction for improving the practical reliability of synthetic augmentation in low-data regimes. The framework is implemented as a modular CLI-based pipeline, where generation, filtering, selection, and validation components can be independently configured and replaced. This design emphasizes reproducibility, flexibility, and practical deployment in real-world data workflows. Through empirical evaluation centered on human pose estimation, we show that synthetic data improves a real-data baseline when used as near-zero-human-annotation-cost augmentation alongside real anchors, while synthetic-only training remains substantially below real-only performance. Supplementary segmentation diagnostics show the same domain-gap pattern. These results highlight the practical value of data-centric orchestration for low-data augmentation.

数据增强合成数据低数据扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。