构建首个孟加拉语书籍推荐大规模异构图数据集,推动低资源语言个性化推荐研究。
Towards Personalized Bangla Book Recommendation: A Large-Scale Heterogeneous Book Graph Dataset
- 构建包含百万级关系的异构书籍图谱,整合作者、用户、评论等多维度信息。
- 实验证明异构关系与混合文本特征对推荐效果影响显著,优于传统方法。
- 适合从事低资源语言推荐、文化语境下个性化系统研究者使用。
孟加拉语文学的个性化书籍推荐受限于缺乏结构化、大规模且公开的数据集。本文提出 RokomariBG,一个用于低资源语言环境下个性化推荐研究的大规模异构书籍图数据集。该数据集包含127,302本书、63,723名用户、16,601位作者、1,515个类别、2,757家出版社和209,602条评论,通过多种关系类型连接,构成综合性知识图谱。为验证数据集价值,我们系统性地在Top-N与序列推荐任务上评估了多种代表性推荐模型。综合基准测试表明,该领域推荐性能受异构关系信息与代码混合文本元数据双重显著影响。这些发现揭示了孟加拉电商生态中尚未被现有推荐基准涵盖的独特挑战。本工作建立了一个基础性基准与公开可用资源,支持可复现评估及未来在低资源文化领域的推荐研究。数据集与代码已公开于 https://github.com/backlashblitz/Bangla-Book-Recommendation-Dataset。
原文摘要 · Abstract (English)
Personalized book recommendation in Bangla literature has been constrained by the lack of structured, large-scale, and publicly available datasets. This work introduces RokomariBG, a large-scale heterogeneous book graph dataset designed to support research on personalized recommendation in a low-resource language setting. The dataset comprises 127,302 books, 63,723 users, 16,601 authors, 1,515 categories, 2,757 publishers, and 209,602 reviews, connected through several relation types and organized as a comprehensive knowledge graph. To demonstrate the utility of the dataset, we present a systematic benchmarking study on the top-N recommendation and sequential recommendation tasks, evaluating a diverse set of representative recommendation models. Through comprehensive benchmarking, we demonstrate that recommendation performance in this domain is strongly influenced by both heterogeneous relational information and code-mixed textual metadata. These findings reveal unique challenges of Bangladeshi e-commerce ecosystems that are largely absent from existing recommendation benchmarks. Overall, this work establishes a foundational benchmark and a publicly available resource for Bangla book recommendation research, enabling reproducible evaluation and future studies on recommendation in low-resource cultural domains. The dataset and code are publicly available at https://github.com/backlashblitz/Bangla-Book-Recommendation-Dataset
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。