arXiv:2511.03020cs.CRcs.AI2025-11

分析电商攻击模式,发现节日期间攻击更严重且可预测。

Exploratory Analysis of Cyberattack Patterns on E-Commerce Platforms Using Statistical Methods

  • 用统计与机器学习结合方法分析攻击趋势。
  • 节日期间攻击更频繁,涉及个人隐私数据威胁更高。
  • 模型可提前预警,适合安全团队用于资源调度。

电商平台网络攻击日益复杂,威胁用户信任与运营稳定。本研究提出一种融合统计建模与机器学习的混合分析框架,基于Verizon社区数据泄露(VCDB)数据集,采用Auto ARIMA进行时间序列预测,并通过曼-惠特尼U检验(U = 2579981.5,p = 0.0121)证实:节日期间攻击严重程度显著高于非节日期间。方差分析(ANOVA)用于考察威胁严重性的季节性差异,集成学习模型(XGBoost、LightGBM、CatBoost)用于预测分类。结果显示,黑色星期五及节日期间存在攻击高峰,涉及个人身份信息(PII)的泄露表现出更高的威胁指标。其中CatBoost表现最佳(准确率85.29%,F1分数0.2254,ROC AUC 0.8247)。该框架首次将季节性预测与可解释的集成学习结合,实现风险时段预判与攻击类型分类。研究考虑了敏感数据使用伦理与偏差评估。尽管存在类别不平衡和依赖历史数据的问题,仍为前瞻性网络安全资源配置提供依据,并指明实时威胁检测的未来方向。

原文摘要 · Abstract (English)

Cyberattacks on e-commerce platforms have grown in sophistication, threatening consumer trust and operational continuity. This research presents a hybrid analytical framework that integrates statistical modelling and machine learning for detecting and forecasting cyberattack patterns in the e-commerce domain. Using the Verizon Community Data Breach (VCDB) dataset, the study applies Auto ARIMA for temporal forecasting and significance testing, including a Mann-Whitney U test (U = 2579981.5, p = 0.0121), which confirmed that holiday shopping events experienced significantly more severe cyberattacks than non-holiday periods. ANOVA was also used to examine seasonal variation in threat severity, while ensemble machine learning models (XGBoost, LightGBM, and CatBoost) were employed for predictive classification. Results reveal recurrent attack spikes during high-risk periods such as Black Friday and holiday seasons, with breaches involving Personally Identifiable Information (PII) exhibiting elevated threat indicators. Among the models, CatBoost achieved the highest performance (accuracy = 85.29%, F1 score = 0.2254, ROC AUC = 0.8247). The framework uniquely combines seasonal forecasting with interpretable ensemble learning, enabling temporal risk anticipation and breach-type classification. Ethical considerations, including responsible use of sensitive data and bias assessment, were incorporated. Despite class imbalance and reliance on historical data, the study provides insights for proactive cybersecurity resource allocation and outlines directions for future real-time threat detection research.

网络安全攻击模式机器学习数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。