通过细粒度引用分析,实现科学文献的自适应嵌入表示
FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
- 基于引用意图与频率设计三元组采样,增强结构信号
- 按背景/方法/结果划分文档细粒度片段,提升领域泛化能力
- 无需微调即可自适应融合多片段表示,适合跨任务应用
科学文献表示学习为多种任务提供强大嵌入,但现有方法在三方面存在挑战:1)对比训练中引用结构信号利用不足,仍生成单一向量表示;2)细粒度表示(句或方面级)需复杂整合且缺乏领域泛化性;3)任务感知学习依赖人工预定义任务分类,忽略细微差异并需额外训练数据。为此,我们提出统一三类方法的新框架FLeW。引入新型三元组采样策略,利用引用意图(背景、方法、结果)与频率增强引用结构信号。引用意图与科学写作通用结构对齐,支持领域泛化的细粒度片段划分。随后采用简单权重搜索,自适应融合三个细粒度嵌入生成任务特定文档表示,无需任务感知微调。实验表明,FLeW在多个科学任务与领域中均展现出良好适用性与鲁棒性,优于先前模型。
原文摘要 · Abstract (English)
Scientific document representation learning provides powerful embeddings for various tasks, while current methods face challenges across three approaches. 1) Contrastive training with citation-structural signals underutilizes citation information and still generates single-vector representations. 2) Fine-grained representation learning, which generates multiple vectors at the sentence or aspect level, requires costly integration and lacks domain generalization. 3) Task-aware learning depends on manually predefined task categorization, overlooking nuanced task distinctions and requiring extra training data for task-specific modules. To address these problems, we propose a new method that unifies the three approaches for better representations, namely FLeW. Specifically, we introduce a novel triplet sampling method that leverages citation intent and frequency to enhance citation-structural signals for training. Citation intents (background, method, result), aligned with the general structure of scientific writing, facilitate a domain-generalized facet partition for fine-grained representation learning. Then, we adopt a simple weight search to adaptively integrate three facet-level embeddings into a task-specific document embedding without task-aware fine-tuning. Experiments show the applicability and robustness of FLeW across multiple scientific tasks and fields, compared to prior models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。