用文本生成合成数据,让模型从零开始高效训练。
GALOT: Generative Active Learning via Optimizable Zero-shot Text-to-image Generation
- 用主动学习优化文本提示,生成更丰富有效的图像数据。
- 仅靠文本描述就能训练出性能优于传统方法的模型。
- 适合数据稀缺场景,降低标注成本,适合视觉模型快速构建。
主动学习(AL)是机器学习中的关键方法,通过识别并利用最具信息量的样本实现高效模型训练。然而,传统主动学习严重依赖有限的标注数据和数据分布,限制了其性能。本文提出GALOT框架,将零样本文本到图像(T2I)生成与主动学习结合,仅通过文本描述即可高效训练机器学习模型。具体而言,利用主动学习准则优化文本输入,生成更具信息量和多样性的合成图像,并基于文本构建伪标签形成合成数据集用于主动学习。该方法降低了数据采集与标注成本,提升了模型训练效率,实现了从文本描述到视觉模型的端到端新范式。在多项实验中,本框架持续显著优于传统主动学习方法。
原文摘要 · Abstract (English)
Active Learning (AL) represents a crucial methodology within machine learning, emphasizing the identification and utilization of the most informative samples for efficient model training. However, a significant challenge of AL is its dependence on the limited labeled data samples and data distribution, resulting in limited performance. To address this limitation, this paper integrates the zero-shot text-to-image (T2I) synthesis and active learning by designing a novel framework that can efficiently train a machine learning (ML) model sorely using the text description. Specifically, we leverage the AL criteria to optimize the text inputs for generating more informative and diverse data samples, annotated by the pseudo-label crafted from text, then served as a synthetic dataset for active learning. This approach reduces the cost of data collection and annotation while increasing the efficiency of model training by providing informative training samples, enabling a novel end-to-end ML task from text description to vision models. Through comprehensive evaluations, our framework demonstrates consistent and significant improvements over traditional AL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。