用论文内部结构做无引用预训练,提升科学文献表征效果
Asymmetric Within-Document Predictive Learning for Scientific Document Representation

- 通过标题/摘要预测方法、方法预测结论实现不对称预测
- 引入正则化后性能显著提升,接近对比学习基线
- 不同编码分支适配不同检索场景,可灵活应用
我们研究利用论文话语结构进行科学文档表征的预测式预训练。提出 SciJEPA 框架,一种无引用的方法:用标题和摘要表示来预测方法部分表示,再用方法表示预测结论表示。在 RELISH、高影响力引文、SciDocs 及引文预测任务上,纯预测训练虽可行但弱于使用相同段落对的受控对比基线。加入分片各向同性高斯正则化(SIGReg)后性能显著提升,缩小了差距。正则化效果具有任务依赖性:中等程度的 SIGReg 有助于细粒度排序,而过强则削弱局部对齐。进一步表明不同编码分支支持不同检索模式。这些结果表明,在嵌入几何受控的前提下,文档内预测学习是科学文档表征的一种有前景的无引用补充方案。
原文摘要 · Abstract (English)
We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using the same section pairs. Adding Sliced Isotropic Gaussian Regularization (SIGReg) substantially improves performance and narrows this gap. The effect of regularization is task-dependent: moderate SIGReg helps fine-grained ranking, while stronger regularization can weaken local alignment. We further show that different encoding branches support different retrieval regimes. These results position within-document predictive learning as a promising citation-free complement for scientific document representation, provided that embedding geometry is carefully controlled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。