流式持续学习的评估结果受时间分段方式影响,同一数据流不同切分会得出不同结论。
Temporal Taskification in Streaming Continual Learning: A Source of Evaluation Instability

- 通过塑性-稳定性分析框架,量化不同时间切分对学习模式的影响。
- 9、30、44天切分下预测误差与遗忘率差异显著,证明分段方式直接影响结果。
- 短切分更敏感,适合研究评估鲁棒性或边界效应的学者参考。
流式持续学习通常通过时间分段将连续数据流转化为离散任务序列。本文认为,这种时间分段并非中立预处理,而是评估体系的结构性组成部分:同一数据流的不同有效切分可能引出不同的持续学习范式,从而导致不同的基准结论。为此,我们提出基于塑性与稳定性轮廓的分段级分析框架,引入轮廓距离与边界-轮廓敏感性(BPS),诊断在训练前小边界扰动如何改变学习范式。在网络流量预测任务中,使用CESNET-Timeseries24数据集,固定数据流、模型和训练预算,仅改变时间分段方式,结果表明在9、30、44天切分下,预测误差、遗忘率和反向迁移均出现显著变化,证明分段本身即可显著影响评估结果。此外,较短切分产生更嘈杂的分布模式、更大的结构距离和更高的BPS,说明其对边界扰动更敏感。这些发现表明,流式持续学习的基准结论不仅取决于学习器和数据,还取决于如何分段,呼吁将时间分段作为评估的首要变量。
原文摘要 · Abstract (English)
Streaming Continual Learning (CL) typically converts a continuous stream into a sequence of discrete tasks through temporal partitioning. We argue that this temporal taskification step is not a neutral preprocessing choice, but a structural component of evaluation: different valid splits of the same stream can induce different CL regimes and therefore different benchmark conclusions. To study this effect, we introduce a taskification-level framework based on plasticity and stability profiles, a profile distance between taskifications, and Boundary-Profile Sensitivity (BPS), which diagnoses how strongly small boundary perturbations alter the induced regime before any CL model is trained. We evaluate continual finetuning, Experience Replay, Elastic Weight Consolidation, and Learning without Forgetting on network traffic forecasting with CESNET-Timeseries24, keeping the stream, model, and training budget fixed while varying only the temporal taskification. Across 9-, 30-, and 44-day splits, we observe substantial changes in forecasting error, forgetting, and backward transfer, showing that taskification alone can materially affect CL evaluation. We further find that shorter taskifications induce noisier distribution-level patterns, larger structural distances, and higher BPS, indicating greater sensitivity to boundary perturbations. These results show that benchmark conclusions in streaming CL depend not only on the learner and the data stream, but also on how that stream is taskified, motivating temporal taskification as a first-class evaluation variable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。