用生成方法提升网络传输慢速预测准确率,发现简单采样更有效
Improving Slow Transfer Predictions: Generative Methods Compared
- 对比传统过采样与生成模型(如CTGAN)的增广策略
- 在极端不平衡数据下,生成方法性能未显著优于简单分层采样
- 适合关注网络性能预测与数据不平衡问题的研究者
科学计算网络中的数据传输性能监控至关重要。通过在通信初期预测性能,可识别潜在慢速传输并进行选择性监控,从而优化网络使用和整体性能。机器学习模型在此任务中面临的关键瓶颈是类别不平衡问题。本研究聚焦于解决该问题以提升预测准确性。我们分析并比较了多种增广策略,包括传统过采样方法与生成技术,并调整训练数据集中的类别不平衡比率,评估其对模型性能的影响。尽管增广可能带来性能提升,但随着不平衡比率增大,性能改善并不显著。结论表明,即使最先进的技术如CTGAN,也未能明显优于简单的分层采样方法。
原文摘要 · Abstract (English)
Monitoring data transfer performance is a crucial task in scientific computing networks. By predicting performance early in the communication phase, potentially sluggish transfers can be identified and selectively monitored, optimizing network usage and overall performance. A key bottleneck to improving the predictive power of machine learning (ML) models in this context is the issue of class imbalance. This project focuses on addressing the class imbalance problem to enhance the accuracy of performance predictions. In this study, we analyze and compare various augmentation strategies, including traditional oversampling methods and generative techniques. Additionally, we adjust the class imbalance ratios in training datasets to evaluate their impact on model performance. While augmentation may improve performance, as the imbalance ratio increases, the performance does not significantly improve. We conclude that even the most advanced technique, such as CTGAN, does not significantly improve over simple stratified sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。