arXiv:2501.00277stat.MLcs.AI2025-01被引 1

提出新框架,让专家高效标注数据,提升模型精度

Efficient Human-in-the-Loop Active Learning: A Novel Framework for Data Labeling in AI Systems

  • 融合多种提问方式,自动决定下一步提问策略
  • 在5个真实数据集上,准确率更高,损失更低
  • 适合医疗等需专家标注的高成本场景

现代AI算法依赖标注数据,但现实中多数数据未标注,标注成本高昂,尤其在医学影像等专业领域。为高效利用专家时间,本文提出一种新型人机协同主动学习框架。不同于传统方法仅关注选择标注样本,本框架创新性地整合不同查询方式的信息,并建立模型以自动决定下一问题的提问方式。进一步引入数据驱动的探索与利用机制,可嵌入多种主动学习算法。在五个真实数据集(包括一个复杂图像任务)上的仿真结果显示,该框架在准确率和损失方面均优于现有方法。

原文摘要 · Abstract (English)

Modern AI algorithms require labeled data. In real world, majority of data are unlabeled. Labeling the data are costly. this is particularly true for some areas requiring special skills, such as reading radiology images by physicians. To most efficiently use expert's time for the data labeling, one promising approach is human-in-the-loop active learning algorithm. In this work, we propose a novel active learning framework with significant potential for application in modern AI systems. Unlike the traditional active learning methods, which only focus on determining which data point should be labeled, our framework also introduces an innovative perspective on incorporating different query scheme. We propose a model to integrate the information from different types of queries. Based on this model, our active learning frame can automatically determine how the next question is queried. We further developed a data driven exploration and exploitation framework into our active learning method. This method can be embedded in numerous active learning algorithms. Through simulations on five real-world datasets, including a highly complex real image task, our proposed active learning framework exhibits higher accuracy and lower loss compared to other methods.

主动学习人机协同数据标注高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。