arXiv:2503.09221cs.CV2025-03被引 4

用主动学习指导生成最有益的分割数据,提升模型性能。

Active Learning Inspired ControlNet Guidance for Augmenting Semantic Segmentation Datasets

  • 在扩散过程反向阶段引入主动学习指标引导生成
  • 训练免费,仅修改反向扩散流程,适配任意预训练ControlNet
  • 生成数据使分割模型性能优于普通合成数据

基于扩散模型的条件图像生成近年展现出优异的图像质量与用户约束保持能力。尤其ControlNet可实现真实分割掩码与生成内容的精确对齐,从而增强分割任务的训练数据。这引发关键问题:能否让ControlNet进一步生成对特定任务最有信息量的合成样本?受主动学习启发,即根据样本难易度或模型不确定性选择最具信息量的真实样本,我们首次将主动学习中的选择度量融入反向扩散过程以指导样本生成。具体探索了不确定性、委员会查询和预期模型变化三种常用主动学习指标,并通过梯度近似实现其在生成过程中的应用。该方法无需训练,仅修改反向扩散流程,可应用于任何预训练ControlNet。实验表明,使用引导生成的合成数据训练的分割模型性能超越使用非引导合成数据的模型。本工作强调了扩散模型需具备不仅对齐图像内容,还需提升下游任务表现的高级控制机制,凸显合成数据生成的真正潜力。

原文摘要 · Abstract (English)

Recent advances in conditional image generation from diffusion models have shown great potential in achieving impressive image quality while preserving the constraints introduced by the user. In particular, ControlNet enables precise alignment between ground truth segmentation masks and the generated image content, allowing the enhancement of training datasets in segmentation tasks. This raises a key question: Can ControlNet additionally be guided to generate the most informative synthetic samples for a specific task? Inspired by active learning, where the most informative real-world samples are selected based on sample difficulty or model uncertainty, we propose the first approach to integrate active learning-based selection metrics into the backward diffusion process for sample generation. Specifically, we explore uncertainty, query by committee, and expected model change, which are commonly used in active learning, and demonstrate their application for guiding the sample generation process through gradient approximation. Our method is training-free, modifying only the backward diffusion process, allowing it to be used on any pretrained ControlNet. Using this process, we show that segmentation models trained with guided synthetic data outperform those trained on non-guided synthetic data. Our work underscores the need for advanced control mechanisms for diffusion-based models, which are not only aligned with image content but additionally downstream task performance, highlighting the true potential of synthetic data generation.

主动学习图像生成数据增强分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。