arXiv:2412.09118cs.LG2024-12

提出算法视角的流数据建模方法,解决概念漂移难题。

An Algorithm-Centered Approach To Model Streaming Data

  • 从算法内部机制出发,用窗口化方式建模流数据
  • 理论分析表明该框架与现有模型在多数场景等价
  • 适用于关键基础设施等实时性要求高的场景

除了经典的离线机器学习设置外,流学习是一种数据随时间持续到达、可能处于非平稳环境中的成熟设置。概念漂移——底层分布随时间变化的现象——构成了重大挑战。尽管具有高度实际意义,但目前缺乏与离线设置中经典统计学习理论相媲美的基础理论。这可归因于缺乏类似概率分布的底层对象。虽然已有方法尝试将思想迁移至流设置,但这些方法均从数据角度出发,而非算法视角。本文提出一种面向算法视角的时序数据新模型。不同于以时间点定义设置,我们采用窗口化方法,模拟大多数流学习算法的内部运作。我们在理论上对比了该框架与其他文献模型,发现多数情况下两者描述的是同一情形。此外,我们进行了数值评估,并展示了其在关键基础设施领域的应用。

原文摘要 · Abstract (English)

Besides the classical offline setup of machine learning, stream learning constitutes a well-established setup where data arrives over time in potentially non-stationary environments. Concept drift, the phenomenon that the underlying distribution changes over time poses a significant challenge. Yet, despite high practical relevance, there is little to no foundational theory for learning in the drifting setup comparable to classical statistical learning theory in the offline setting. This can be attributed to the lack of an underlying object comparable to a probability distribution as in the classical setup. While there exist approaches to transfer ideas to the streaming setup, these start from a data perspective rather than an algorithmic one. In this work, we suggest a new model of data over time that is aimed at the algorithm's perspective. Instead of defining the setup using time points, we utilize a window-based approach that resembles the inner workings of most stream learning algorithms. We compare our framework to others from the literature on a theoretical basis, showing that in many cases both model the same situation. Furthermore, we perform a numerical evaluation and showcase an application in the domain of critical infrastructure.

流学习概念漂移算法视角窗口模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。