让科学论文嵌入可拆解、可解释,支持精准语义分析与可视化。
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
- 通过无监督方法生成针对不同科学维度的摘要句,构建多维度嵌入。
- 单次前向传播即可输出多个维度的嵌入,效率高且支持可控相似性计算。
- 可将嵌入还原为自然语言描述,即使在空白区域也保持可解释性。
我们提出 SemCSE-Multi,一种用于生成科学摘要多维度嵌入的新型无监督框架,评估领域涵盖入侵生物学与医学。该嵌入能独立捕捉并指定不同方面,实现细粒度、可控制的相似性评估及用户驱动的科学领域自适应可视化。方法基于无监督流程生成特定于方面的摘要句,并训练嵌入模型将语义相关的摘要映射到嵌入空间中相近位置。随后,我们将这些方面特定的嵌入能力提炼为统一模型,可在一次前向传播中直接从科学摘要预测多个方面嵌入。此外,我们引入嵌入解码管道,可将嵌入还原为对应方面的自然语言描述。值得注意的是,该解码在低维可视化中的未占据区域依然有效,显著提升以用户为中心场景下的可解释性。
原文摘要 · Abstract (English)
We propose SemCSE-Multi, a novel unsupervised framework for generating multifaceted embeddings of scientific abstracts, evaluated in the domains of invasion biology and medicine. These embeddings capture distinct, individually specifiable aspects in isolation, thus enabling fine-grained and controllable similarity assessments as well as adaptive, user-driven visualizations of scientific domains. Our approach relies on an unsupervised procedure that produces aspect-specific summarizing sentences and trains embedding models to map semantically related summaries to nearby positions in the embedding space. We then distill these aspect-specific embedding capabilities into a unified embedding model that directly predicts multiple aspect embeddings from a scientific abstract in a single, efficient forward pass. In addition, we introduce an embedding decoding pipeline that decodes embeddings back into natural language descriptions of their associated aspects. Notably, we show that this decoding remains effective even for unoccupied regions in low-dimensional visualizations, thus offering vastly improved interpretability in user-centric settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。