基于预测目标的智能采样,提升流数据学习效率
Prediction-Oriented Subsampling from Data Streams
- 以降低下游预测不确定性为目标进行数据采样
- 在两个基准任务上优于已有信息论方法
- 适合追求高效流数据训练的研究者
数据常以流的形式持续生成,学习模型面临如何在控制计算成本的前提下捕捉有效信息的挑战。本文研究离线学习中的智能数据采样,提出一种以减少关键下游预测不确定性为核心的信息论方法。实验表明,该预测导向的方法在两个广泛研究的问题上表现优于先前的信息论技术。同时强调,实际中取得良好性能需精心设计模型架构。
原文摘要 · Abstract (English)
Data is often generated in streams, with new observations arriving over time. A key challenge for learning models from data streams is capturing relevant information while keeping computational costs manageable. We explore intelligent data subsampling for offline learning, and argue for an information-theoretic method centred on reducing uncertainty in downstream predictions of interest. Empirically, we demonstrate that this prediction-oriented approach performs better than a previously proposed information-theoretic technique on two widely studied problems. At the same time, we highlight that reliably achieving strong performance in practice requires careful model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。