arXiv:2511.20605cs.LGstat.ML2025-11中稿 · publication in INF…被引 1

用主动学习市场低价买标签,省钱又高效

How to Purchase Labels? A Cost-Effective Approach Using Active Learning Markets

  • 设计买卖双方市场,用优化算法决定买哪些标签
  • 少用标签就达到更好模型效果,实测提升显著
  • 适合预算有限、想高效标注数据的从业者

我们提出并分析了主动学习市场作为一种购买标签的新方式,适用于分析师希望获取额外数据以改进模型拟合或训练预测模型的情形。与现有大量购买特征和样本的方案不同,本文首次将市场清算法形式化为优化问题,整合预算约束和性能提升阈值。聚焦单买家多卖家场景,采用基于方差和基于委员会查询的两种主动学习策略,并搭配不同定价机制,与随机采样及贪婪背包启发式等基线方法对比。在房地产估价和能源预测两个真实世界数据集上验证,结果表明该方法鲁棒性强,以更少标签获得更优性能。本方案提供了一种易实现的资源受限环境下的数据采集优化工具。

原文摘要 · Abstract (English)

We introduce and analyse active learning markets as a way to purchase labels, in situations where analysts aim to acquire additional data to improve model fitting, or to better train models for predictive analytics applications. This comes in contrast to the many proposals that already exist to purchase features and examples. By originally formalising the market clearing as an optimisation problem, we integrate budget constraints and improvement thresholds into the label acquisition process. We focus on a single-buyer-multiple-seller setup and propose the use of two active learning strategies (variance based and query-by-committee based), paired with distinct pricing mechanisms. They are compared to benchmark baselines including random sampling and a greedy knapsack heuristic. The proposed strategies are validated on real-world datasets from two critical application domains: real estate pricing and energy forecasting. Results demonstrate the robustness of our approach, consistently achieving superior performance with fewer labels acquired compared to conventional methods. Our proposal comprises an easy-to-implement practical solution for optimising data acquisition in resource-constrained environments.

主动学习数据标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。