用大模型做主动学习,既能选数据又能生成新数据,提升效率。
From Selection to Generation: A Survey of LLM-based Active Learning
- 用大模型自动筛选最有价值的数据,减少人工标注
- 大模型可生成高质量新数据,降低标注成本
- 适合想高效训练模型的研究者和开发者
主动学习(AL)通过选择最具信息量的数据进行标注和训练,显著提升模型效率与性能。近年来,大语言模型(LLMs)不仅用于数据选择,还能生成全新数据实例并提供更低成本的标注。随着大模型时代对高质量数据和高效训练需求的增长,本文系统综述了基于大模型的主动学习技术。提出一个直观的分类体系,梳理大模型在主动学习流程中的多重角色。分析主动学习对大模型学习范式的影响,并探讨其在多领域的应用。最后指出当前挑战并展望未来研究方向。本综述旨在为研究人员和实践者提供关于大模型主动学习技术的清晰理解,助力其在新场景中部署应用。
原文摘要 · Abstract (English)
Active Learning (AL) has been a powerful paradigm for improving model efficiency and performance by selecting the most informative data points for labeling and training. In recent active learning frameworks, Large Language Models (LLMs) have been employed not only for selection but also for generating entirely new data instances and providing more cost-effective annotations. Motivated by the increasing importance of high-quality data and efficient model training in the era of LLMs, we present a comprehensive survey on LLM-based Active Learning. We introduce an intuitive taxonomy that categorizes these techniques and discuss the transformative roles LLMs can play in the active learning loop. We further examine the impact of AL on LLM learning paradigms and its applications across various domains. Finally, we identify open challenges and propose future research directions. This survey aims to serve as an up-to-date resource for researchers and practitioners seeking to gain an intuitive understanding of LLM-based AL techniques and deploy them to new applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。