arXiv:2501.04364cs.IR2025-01被引 6

用新方法直接采集结构化用户数据,省去繁琐日志处理

An innovative data collection method to eliminate the preprocessing phase in web usage mining

  • 通过应用级API直接收集结构化日志,跳过传统服务器日志处理
  • 采集数据可直接用于实时分析与推荐系统,性能更优
  • 适合需要高效用户行为分析的电商平台与企业应用

网页使用挖掘(WUM)通常依赖服务器日志作为数据源,但此类日志对客户端信息记录有限,且需大量工作才能识别会话,结果常不理想,难以高效用于网页分析。本文提出一种创新的用户追踪与会话管理方法,基于新型数据采集策略,通过应用级API获取并处理日志数据。该技术已成功集成至企业级网络应用中,所收集的同质化结构化数据存储于关系型数据库,相比服务器日志更易于浏览、筛选与处理,可直接作为高性能网页使用挖掘、实时网页分析或推荐系统的核心数据源。

原文摘要 · Abstract (English)

The underlying data source for web usage mining (WUM) is commonly thought to be server logs. However, access log files ensure quite limited data about the clients. Identifying sessions from this messy data takes a considerable effort, and operations performed for this purpose do not always yield excellent results. Also, this data cannot be used for web analytics efficiently. This study proposes an innovative method for user tracking, session management, and collecting web usage data. The method is mainly based on a new approach for using collected data for web analytics extraction as the data source in web usage mining. An application-based API has been developed with a different strategy from conventional client-side methods to obtain and process log data. The log data has been successfully gathered by integrating the technique into an enterprise web application. The results reveal that the homogeneous structured data collected and stored with this method is more convenient to browse, filter, and process than web server logs. This data stored on a relational database can be used effortlessly as a reliable data source for high-performance web usage mining activity, real-time web analytics, or a functional recommendation system.

用户行为分析数据采集实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。