arXiv:2511.03877cs.LG2025-11KDD

构建社交平台长时滞预测基准数据集,助力早期行为预测长期影响。

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

  • 提出'先行-滞后预测'新范式,用早期互动预测后期高影响力结果。
  • 发布arXiv(230万论文)与GitHub(300万仓库)两大基准数据集,覆盖多年动态。
  • 适合研究社交网络、用户行为预测及长周期时间序列建模的学者。

社交与协作平台产生多变量时间序列数据,其中早期互动(如浏览、点赞、下载)可能在数月甚至数年后才引发高影响力结果(如引用、销售、评论)。本文将此现象形式化为先行-滞后预测(LLF):给定一个早期使用通道(先行),预测一个时序错位但相关的后续结果通道(滞后)。尽管此类模式普遍存在,但因缺乏标准化数据集,该问题尚未被时间序列领域统一研究。为此,本文构建了两个高容量基准数据集:arXiv(230万篇论文的访问量 → 引用关系)和GitHub(300万仓库的推送/星标 → 分支数)。数据集涵盖跨年长时滞动态,覆盖全范围结果,且采样避免幸存者偏差。我们详细记录数据清洗过程,通过统计与分类测试验证了先行-滞后关系的存在,并对参数与非参数基线模型进行了回归性能基准测试。本研究确立了LLF作为新型预测范式,并为其在社交与使用数据中的系统性探索奠定实证基础。

原文摘要 · Abstract (English)

Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years later, by higher impact like citations, sales, or reviews. We formalize this setting as Lead-Lag Forecasting (LLF): given an early usage channel (the lead), predict a correlated but temporally shifted outcome channel (the lag). Despite the ubiquity of such patterns, LLF has not been treated as a unified forecasting problem within the time-series community, largely due to the absence of standardised datasets. To anchor research in LLF, here we present two high-volume benchmark datasets: arXiv (accesses -> citations of 2.3M papers) and GitHub (pushes/stars -> forks of 3M repositories). Our datasets provide ideal testbeds for lead-lag forecasting, by capturing long-horizon dynamics across years, spanning the full spectrum of outcomes, and avoiding survivorship bias in sampling. We documented all technical details of data curation and cleaning, verified the presence of lead-lag dynamics through statistical and classification tests, and benchmarked parametric and non-parametric baselines for regression. Our study establishes LLF as a novel forecasting paradigm and lays an empirical foundation for its systematic exploration in social and usage data.

时间序列预测建模社交网络数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。