arXiv:2604.16329cs.IRcs.AI2026-04

让论文推荐可控制多样性,按背景和方法分开打分。

Beyond Single-Score Ranking: Facet-Aware Reranking for Controllable Diversity in Paper Recommendation

论文配图:Beyond Single-Score Ranking: Facet-Aware Reranking for Controllable Diversity in Paper Recommendation
图 1 · 摘自论文原文
  • 拆分论文相关性为背景与方法两个独立维度建模
  • 背景和方法的推荐效果分别提升5.9和31.1点NDCG@20
  • 用5891条真实标注数据比4万条合成数据更高效

现有论文推荐系统仅输出单一相似度分数,混杂了多种相关性含义,用户无法指定相似的原因。我们提出SciFACE(Scientific Faceted Cross-Encoder),一个对背景(研究问题)和方法(解决方式)两个独立维度进行建模的重排序框架。SciFACE在5,891对由GPT-4o-mini标注的种子-候选论文对上训练两个独立交叉编码器,标签依据特定维度标准生成,并经人工判断验证。在CSFCube数据集上,背景维度达到70.63 NDCG@20(比SPECTER高5.9点),方法维度达49.06 NDCG@20(比SPECTER高31.1点),性能接近当前最优水平。相比无引文预训练的FaBLE,SciFACE以5,891条真实标注数据实现方法维度4.1点的提升,显著优于使用4万条合成增强数据的方案。结果表明,高质量的有根基的细粒度标签在学习科学相似性时比大规模合成数据更具数据效率。

原文摘要 · Abstract (English)

Current paper recommendation systems output a single similarity score that mixes different notions of relatedness, so users cannot specify why papers should be similar. We present SciFACE (Scientific Faceted Cross-Encoder), a reranking framework that models two independent facets: Background (what problem is studied) and Method (how it is solved). SciFACE trains two separate cross-encoders on 5,891 real seed-candidate paper pairs labeled by GPT-4o-mini with facet-specific criteria and validated against human judgments. On CSFCube, SciFACE reaches 70.63 NDCG@20 on Background (5.9 points above SPECTER) and 49.06 NDCG@20 on Method (31.1 points above SPECTER), competitive with state-of-the-art results. Compared with FaBLE without citation pre-training, SciFACE improves Method NDCG@20 by 4.1 points while using 5,891 labeled pairs versus 40K synthetic augmentations. These results show that high-quality grounded facet labels can be more data-efficient than large-scale synthetic augmentation for learning fine-grained scientific similarity.

论文推荐多维排序细粒度匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。