用多源日志融合提升网页行为预测与异常检测精度
Predictive modeling and anomaly detection in large-scale web portals through the CAWAL framework
- 融合应用日志与网页分析数据,构建更完整的用户会话与页面视图数据集
- 在大型网页门户上实现超92%的用户行为预测准确率
- 适合需要高可靠性与可扩展性的大规模网站运维团队
本研究提出一种基于CAWAL框架的方法,通过整合应用日志与网页分析数据,生成丰富的会话与页面浏览数据集,用于高级预测建模与异常检测。传统网页使用挖掘(WUM)依赖服务器日志,数据多样性与质量受限。该框架通过跨集成会话与页面视图数据,显著提升数据多样性与质量,省去传统预处理环节,提高流程效率。利用梯度提升与随机森林等机器学习模型,在真实大规模网页门户上实现超过92%的用户行为预测准确率,并显著增强异常检测能力。结果表明,该方法能提供用户行为与系统性能的深度洞察,为提升大型网页门户的效率、可靠性和可扩展性提供了可靠解决方案。
原文摘要 · Abstract (English)
This study presents an approach that uses session and page view data collected through the CAWAL framework, enriched through specialized processes, for advanced predictive modeling and anomaly detection in web usage mining (WUM) applications. Traditional WUM methods often rely on web server logs, which limit data diversity and quality. Integrating application logs with web analytics, the CAWAL framework creates comprehensive session and page view datasets, providing a more detailed view of user interactions and effectively addressing these limitations. This integration enhances data diversity and quality while eliminating the preprocessing stage required in conventional WUM, leading to greater process efficiency. The enriched datasets, created by cross-integrating session and page view data, were applied to advanced machine learning models, such as Gradient Boosting and Random Forest, which are known for their effectiveness in capturing complex patterns and modeling non-linear relationships. These models achieved over 92% accuracy in predicting user behavior and significantly improved anomaly detection capabilities. The results show that this approach offers detailed insights into user behavior and system performance metrics, making it a reliable solution for improving large-scale web portals' efficiency, reliability, and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。