用大模型嵌入构建科学文献图谱,揭示文本与链接的双重结构。
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
- 融合文本语义与引用关系,构建图文混合的科学文献图谱。
- 基于约5600万篇论文数据,发现文本自组织的内在结构模式。
- 适合研究知识演化、学术影响力或跨领域关联的学者使用。
大规模文本数据集(如学术论文、网页等)包含两类特征:一是文本本身所承载的语义信息;二是通过引用、链接或共享属性与其他文本的关联关系。前者可通过大语言模型嵌入实现高效表征,后者则可建模为图结构并应用经典算法进行分类与预测。本文以包含约5600万篇科学论文的Web of Science数据集为对象,结合所提出的嵌入方法,从图文融合视角分析其结构,揭示出文本自身具备的自组织景观。
原文摘要 · Abstract (English)
Large text data sets, such as publications, websites, and other text-based media, inherit two distinct types of features: (1) the text itself, its information conveyed through semantics, and (2) its relationship to other texts through links, references, or shared attributes. While the latter can be described as a graph structure and can be handled by a range of established algorithms for classification and prediction, the former has recently gained new potential through the use of LLM embedding models. Demonstrating these possibilities and their practicability, we investigate the Web of Science dataset, containing ~56 million scientific publications through the lens of our proposed embedding method, revealing a self-structured landscape of texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。