构建标准化基准,统一评估引文推荐模型性能。
Benchmark for Evaluation and Analysis of Citation Recommendation Models
- 设计统一数据集与评测指标,支持跨模型对比。
- 覆盖引文上下文多维度特征,全面评估模型表现。
- 适合引文推荐研究者和系统开发者使用。
引文推荐系统受到广泛关注,但不同研究在方法、数据集和评估指标上差异显著。有的关注论文整体内容,有的侧重引文上下文;数据集涵盖元信息、引用上下文甚至全文,格式多样。这种多样性导致模型评估与比较困难。为此,我们提出构建一个专门用于分析和比较引文推荐模型的基准。该基准将统一评估模型在引文上下文不同特征上的表现,并以标准化方式呈现结果,为研究人员提供一致的评估平台,促进有效对比,助力发现有前景的研究方向。
原文摘要 · Abstract (English)
Citation recommendation systems have attracted much academic interest, resulting in many studies and implementations. These systems help authors automatically generate proper citations by suggesting relevant references based on the text they have written. However, the methods used in citation recommendation differ across various studies and implementations. Some approaches focus on the overall content of papers, while others consider the context of the citation text. Additionally, the datasets used in these studies include different aspects of papers, such as metadata, citation context, or even the full text of the paper in various formats and structures. The diversity in models, datasets, and evaluation metrics makes it challenging to assess and compare citation recommendation methods effectively. To address this issue, a standardized dataset and evaluation metrics are needed to evaluate these models consistently. Therefore, we propose developing a benchmark specifically designed to analyze and compare citation recommendation models. This benchmark will evaluate the performance of models on different features of the citation context and provide a comprehensive evaluation of the models across all these tasks, presenting the results in a standardized way. By creating a benchmark with standardized evaluation metrics, researchers and practitioners in the field of citation recommendation will have a common platform to assess and compare different models. This will enable meaningful comparisons and help identify promising approaches for further research and development in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。