arXiv:2608.13874cs.IRcs.CY2026-08

预测新发推文在蓝星自定义订阅源中的返回概率。

Predicting Custom-Feed Returns for New Bluesky Posts: A Prospective Study

  • 将新帖子作为查询,排序所有可推荐订阅源的返回可能性。
  • 数据集含1780万条推文、186万次返回记录,测试集9.04%有正样本。
  • 使用LambdaRank模型,在召回率与命中率上表现最优。

传统冷启动推荐针对新用户或新内容。蓝星自定义订阅源则不同:独立运营的订阅源从共享公共流中筛选内容。在此场景下,新发布的帖子是冷启动对象,而订阅源则是候选。本文提出一种冷启动路由任务:以新发布的公共帖子为查询,根据各订阅源后续是否返回该帖进行排序。构建了一个持续收集、延后标注的基准数据集,覆盖5000个监控订阅源,包含1780万条公共帖子、186.5万条可观测的帖子-订阅源返回记录和62.5万次有效订阅源投票。标签定义为:帖子在发布后24小时内至少一次出现在订阅源的AppView Top-50结果中。实验采用两个互斥的24小时测试折,每折配以24小时训练窗口和24小时结果可用间隔。评估基于602,186条至少有一个正标签且满足指标条件的测试帖,占全部666万测试帖的9.04%。在两折中,LambdaRank在各项指标上均表现最佳:召回@10为0.7361,NDCG@10为0.6127,命中率@10为0.7749。

原文摘要 · Abstract (English)

The conventional approach to cold-start recommendation addresses new users or newly introduced items. Bluesky custom feeds create a different setting: independently operated feeds filter content from a shared public stream. In this setting, newly published posts are the cold-start objects, while the feeds serve as candidates. We propose a cold-start routing task in which a newly ingested public post is the query and all rankable feeds in the monitored panel are ranked according to whether each will subsequently return it. We build a still-evolving collect-first, label-later benchmark dataset. The collected dataset covers a fixed panel of 5,000 monitored feeds and contains 17.804 million public posts, 1.865 million observable post--feed return records, and 625,083 valid feed polls. The labels record whether a post is observed among a feed's AppView Top-50 results in at least one poll during the 24 hours after publication. The current experiments use two disjoint 24-hour test folds, each paired with a 24-hour training window and separated by a 24-hour outcome-availability gap. Evaluation is conditional on the 602,186 test posts that have at least one positive observed label and satisfy the metric eligibility criteria; these posts account for 9.04% of all 6,661,658 test posts. Across the two folds, LambdaRank achieves the best equal-fold mean values among the evaluated models: 0.7361 for capped Recall@10, 0.6127 for NDCG@10, and 0.7749 for Hit@10.

冷启动推荐系统社交网络蓝星

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。