arXiv:2602.04567cs.IRcs.CY2026-02中稿 · The ACM Web Confer…被引 8

发布400亿条交互的短视频推荐数据集,助力真实场景算法研究。

VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommendation

  • 构建包含1000万用户、2000万视频的超大规模工业级数据集。
  • 覆盖6个月真实平台行为,含400亿次交互与多维度反馈信号。
  • 适合研究序列推荐、冷启动问题及下一代推荐系统的设计。

短视频推荐面临用户兴趣快速变化等挑战,但受限于缺乏反映真实平台动态的大规模开源数据集。为此,我们推出了目前最大规模的公开工业数据集——VK大型短视频数据集(VK-LSVD)。该数据集涵盖超过400亿条来自1000万用户和近2000万视频在六个月内产生的交互记录,同时包含内容嵌入、多样化反馈信号及上下文元数据等丰富特征。分析表明数据集具有高质量与多样性。其影响力已体现在2025年VK推荐系统挑战赛中。VK-LSVD为构建更真实的基准测试提供了关键支持,可加速序列推荐、冷启动场景及下一代推荐系统的研究。

原文摘要 · Abstract (English)

Short-video recommendation presents unique challenges, such as modeling rapid user interest shifts from implicit feedback, but progress is constrained by a lack of large-scale open datasets that reflect real-world platform dynamics. To bridge this gap, we introduce the VK Large Short-Video Dataset (VK-LSVD), the largest publicly available industrial dataset of its kind. VK-LSVD offers an unprecedented scale of over 40 billion interactions from 10 million users and almost 20 million videos over six months, alongside rich features including content embeddings, diverse feedback signals, and contextual metadata. Our analysis supports the dataset's quality and diversity. The dataset's immediate impact is confirmed by its central role in the live VK RecSys Challenge 2025. VK-LSVD provides a vital, open dataset to use in building realistic benchmarks to accelerate research in sequential recommendation, cold-start scenarios, and next-generation recommender systems.

短视频推荐工业数据集序列推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。