用问答式摘要和多层级对比学习提升论文推荐精准度。
Discourse-Aware Scientific Paper Recommendation via QA-Style Summarization and Multi-Level Contrastive Learning
- 将论文转为问答结构化摘要,增强语义连贯性。
- 在元数据、章节和全文层级做对比学习,提升匹配精度。
- 适合隐私敏感场景下的学术论文推荐系统研究者。
开放获取论文的快速增长加剧了识别相关文献的挑战。由于隐私限制和用户交互数据获取受限,近期研究转向仅依赖文本内容的推荐方法。然而,现有模型通常将论文视为非结构化文本,忽视其论述结构,限制了语义完整性和可解释性。为此,我们提出OMRC-MR,一个分层框架,融合问答式OMRC(目标、方法、结果、结论)摘要、多层级对比学习和结构感知重排序,用于学术推荐。问答摘要模块将原始论文转化为结构化且论述一致的表示,多层级对比目标在元数据、章节和文档层面对齐语义表征,最终重排序阶段通过上下文相似性校准进一步提升检索精度。在DBLP、S2ORC及新构建的Sci-OMRC数据集上的实验表明,OMRC-MR持续优于现有最先进基线,在Precision@10和Recall@10上分别实现最高7.2%和3.8%的提升。额外评估证实,问答式摘要生成更连贯且事实完整的表示。总体而言,OMRC-MR提供了一个统一、可解释的内容驱动范式,推动可信且隐私友好的学术信息检索发展。
原文摘要 · Abstract (English)
The rapid growth of open-access (OA) publications has intensified the challenge of identifying relevant scientific papers. Due to privacy constraints and limited access to user interaction data, recent efforts have shifted toward content-based recommendation, which relies solely on textual information. However, existing models typically treat papers as unstructured text, neglecting their discourse organization and thereby limiting semantic completeness and interpretability. To address these limitations, we propose OMRC-MR, a hierarchical framework that integrates QA-style OMRC (Objective, Method, Result, Conclusion) summarization, multi-level contrastive learning, and structure-aware re-ranking for scholarly recommendation. The QA-style summarization module converts raw papers into structured and discourse-consistent representations, while multi-level contrastive objectives align semantic representations across metadata, section, and document levels. The final re-ranking stage further refines retrieval precision through contextual similarity calibration. Experiments on DBLP, S2ORC, and the newly constructed Sci-OMRC dataset demonstrate that OMRC-MR consistently surpasses state-of-the-art baselines, achieving up to 7.2% and 3.8% improvements in Precision@10 and Recall@10, respectively. Additional evaluations confirm that QA-style summarization produces more coherent and factually complete representations. Overall, OMRC-MR provides a unified and interpretable content-based paradigm for scientific paper recommendation, advancing trustworthy and privacy-aware scholarly information retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。