arXiv:2512.12787cs.LGcs.AI2025-12被引 4

提出多数据集在线回归模型的统计检验方法,验证性能差异是否显著。

Unveiling Statistical Significance of Online Regression over Multiple Datasets

  • 采用弗里德曼检验与事后检验比较多个在线回归模型。
  • 在真实与合成数据上通过5折交叉验证和种子平均评估性能。
  • 发现现有先进方法仍有可提升空间,结果具统计可信度。

尽管大量研究关注单一数据集上两种学习算法的性能评估,但如何在多个数据集间对多个算法进行统计检验这一关键挑战,在多数机器学习研究中仍被忽视。在在线学习领域,确保统计显著性对于验证持续学习过程至关重要,尤其在实现快速收敛及及时应对概念漂移方面。需要稳健的统计方法来评估随时间演化的数据中性能差异的显著性。本文考察了最先进的在线回归模型,并实证评估了几种合适的统计检验方法。为在多个数据集上比较多个在线回归模型,我们采用了弗里德曼检验及其相应的事后检验。为实现全面评估,使用了真实与合成数据,结合5折交叉验证和种子平均。测试结果普遍确认了竞争性基线的性能与各自报告一致。然而,部分统计检验结果也表明,某些先进方法仍存在改进空间。

原文摘要 · Abstract (English)

Despite extensive focus on techniques for evaluating the performance of two learning algorithms on a single dataset, the critical challenge of developing statistical tests to compare multiple algorithms across various datasets has been largely overlooked in most machine learning research. Additionally, in the realm of Online Learning, ensuring statistical significance is essential to validate continuous learning processes, particularly for achieving rapid convergence and effectively managing concept drifts in a timely manner. Robust statistical methods are needed to assess the significance of performance differences as data evolves over time. This article examines the state-of-the-art online regression models and empirically evaluates several suitable tests. To compare multiple online regression models across various datasets, we employed the Friedman test along with corresponding post-hoc tests. For thorough evaluations, utilizing both real and synthetic datasets with 5-fold cross-validation and seed averaging ensures comprehensive assessment across various data subsets. Our tests generally confirmed the performance of competitive baselines as consistent with their individual reports. However, some statistical test results also indicate that there is still room for improvement in certain aspects of state-of-the-art methods.

在线学习统计检验回归模型多数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。