arXiv:2601.15773cs.LG2026-01AAAI被引 1

用多个大模型协作标注,提升自动标注的准确性和稳定性。

Next Generation Active Learning: Mixture of LLMs in the Loop

  • 通过多个大模型协同标注,融合各自优势提高可靠性
  • 在多个数据集上表现接近人工标注,优于单一模型
  • 支持本地运行,适合实际部署场景

随着大语言模型(LLM)的快速发展及其强大的泛化能力,它们被越来越多地引入主动学习流程中作为标注者,以降低标注成本。然而,由于标注质量限制,由LLM生成的标签往往难以满足真实应用需求。为此,我们提出一种新型主动学习框架——基于混合大模型的闭环主动学习(Mixture of LLMs in the Loop Active Learning),用基于多模型集成的标注模型替代人工标注者,通过聚合多个大模型的优势来增强基于大模型的标注鲁棒性。为减轻噪声标签的影响,我们引入标注差异检测与负向学习机制,识别不可靠标注并提升学习效率。大量实验证明,该框架性能接近人工标注水平,并持续优于单个大模型基线及其它基于大模型集成的方法。此外,本框架基于轻量级大模型构建,可在本地设备上完整运行,适用于真实场景应用。

原文摘要 · Abstract (English)

With the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering the annotation quality, labels generated by LLMs often fall short of real-world applicability. To address this, we propose a novel active learning framework, Mixture of LLMs in the Loop Active Learning, replacing human annotators with labels generated through a Mixture-of-LLMs-based annotation model, aimed at enhancing LLM-based annotation robustness by aggregating the strengths of multiple LLMs. To further mitigate the impact of the noisy labels, we introduce annotation discrepancy and negative learning to identify the unreliable annotations and enhance learning effectiveness. Extensive experiments demonstrate that our framework achieves performance comparable to human annotation and consistently outperforms single-LLM baselines and other LLM-ensemble-based approaches. Moreover, our framework is built on lightweight LLMs, enabling it to operate fully on local machines in real-world applications.

主动学习大模型集成标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。