arXiv:2412.16839cs.CVcs.AI2024-12中稿 · TVCG2025被引 14

用人工引导生成图像,扩充小样本数据集

Human-Guided Image Generation for Expanding Small-Scale Training Image Datasets

  • 通过多模态投影实现图像与生成内容的可控探索
  • 用户反馈样本后自动优化提示词,提升生成多样性
  • 适合需要高质量小数据集的视觉任务研究者

某些实际应用(如稀有野生动物观测)中,计算机视觉模型性能受限于可用图像数量少。利用预训练生成模型扩展数据集是有效方法,但自动生成过程不可控,导致生成图像多样性不足且常含不想要的内容。本文提出一种人工引导的图像生成方法,用于更可控的数据集扩展。我们开发了一种具有理论保障的多模态投影方法,便于探索原始与生成图像。基于此探索,用户可优化提示词并重新生成图像以提升性能。由于直接修改提示词对新手困难,我们进一步设计了样本级提示词优化方法:用户仅需提供样本级反馈(如哪些样本不理想),即可获得更优提示词。该方法在多模态投影量化评估、分类与目标检测任务案例研究中的模型性能提升,以及专家正面反馈中得到验证。

原文摘要 · Abstract (English)

The performance of computer vision models in certain real-world applications (e.g., rare wildlife observation) is limited by the small number of available images. Expanding datasets using pre-trained generative models is an effective way to address this limitation. However, since the automatic generation process is uncontrollable, the generated images are usually limited in diversity, and some of them are undesired. In this paper, we propose a human-guided image generation method for more controllable dataset expansion. We develop a multi-modal projection method with theoretical guarantees to facilitate the exploration of both the original and generated images. Based on the exploration, users refine the prompts and re-generate images for better performance. Since directly refining the prompts is challenging for novice users, we develop a sample-level prompt refinement method to make it easier. With this method, users only need to provide sample-level feedback (e.g., which samples are undesired) to obtain better prompts. The effectiveness of our method is demonstrated through the quantitative evaluation of the multi-modal projection method, improved model performance in the case study for both classification and object detection tasks, and positive feedback from the experts.

图像生成数据扩充人机交互小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。