arXiv:2505.17507cs.IR2025-05中稿 · SIGIR 2025被引 3

构建首个开源模型资源知识图谱,支持推荐、分类与溯源任务

Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph

  • 基于Hugging Face构建包含260万节点的领域知识图谱
  • 提出多任务评测基准,支持资源推荐、分类与演化追踪
  • 适用于关注开源模型管理与信息检索的研究者

开源机器学习资源(如模型和数据集)的快速增长推动了信息检索研究。然而,现有平台如Hugging Face缺乏结构化表示,限制了高级查询与分析,如模型演化追踪与相关数据集推荐。为此,我们构建了首个大规模知识图谱HuggingKG,从Hugging Face社区中提取,包含260万节点和620万边,捕捉领域特定关系与丰富文本属性。基于此,我们进一步提出HuggingBench,一个包含三项新测试集的多任务基准,用于信息检索任务中的资源推荐、分类与追踪。实验揭示了HuggingKG及其衍生任务的独特特征。两项资源均公开可用,有望推动开源资源共享与管理研究。

原文摘要 · Abstract (English)

The rapid growth of open source machine learning (ML) resources, such as models and datasets, has accelerated IR research. However, existing platforms like Hugging Face do not explicitly utilize structured representations, limiting advanced queries and analyses such as tracing model evolution and recommending relevant datasets. To fill the gap, we construct HuggingKG, the first large-scale knowledge graph built from the Hugging Face community for ML resource management. With 2.6 million nodes and 6.2 million edges, HuggingKG captures domain-specific relations and rich textual attributes. It enables us to further present HuggingBench, a multi-task benchmark with three novel test collections for IR tasks including resource recommendation, classification, and tracing. Our experiments reveal unique characteristics of HuggingKG and the derived tasks. Both resources are publicly available, expected to advance research in open source resource sharing and management.

知识图谱推荐系统开源生态信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。