用检索增强生成技术构建可信赖的可降解聚合物知识系统
Retrieval Augmented Generation of Literature-derived Polymer Knowledge: The Example of a Biodegradable Polymer Expert System
- 设计向量与图结构双路径检索,保留研究上下文并支持多跳推理
- 图模型在精度和可解释性上更优,向量模型召回率更高,互补性强
- 经领域专家验证,回答有据可查、符合专业判断,适合材料科研人员
聚合物文献包含大量不断增长的实验知识,但大多散落在非结构化文本中,术语不统一,难以系统检索与推理。现有工具多孤立提取特定研究的细小事实,无法保留跨研究的上下文以回答更广泛的科学问题。检索增强生成(RAG)通过结合大语言模型与外部检索有望突破此局限,但其效果高度依赖领域知识的表示方式。本文基于1000余篇聚羟基烷酸酯(PHA)论文,构建了保持上下文的段落嵌入和规范化结构化知识图谱,支持实体消歧与多跳推理。通过标准检索指标、与GPT、Gemini等先进系统的对比,以及领域化学家的定性验证,结果表明:图结构方法(GraphRAG)在精度与可解释性上更优,向量方法(VectorRAG)具有更广召回率,体现互补优势。专家验证确认,定制化管道,尤其是GraphRAG,能生成有依据、可引用、高度相关的可靠回答。该系统通过每条陈述均基于证据,使研究人员能高效导航文献、跨研究比对发现、挖掘人工难提取的规律。本工作为利用精选语料与检索设计构建材料科学助手提供了实用框架,减少对专有模型依赖,实现可信赖的大规模文献分析。
原文摘要 · Abstract (English)
Polymer literature contains a large and growing body of experimental knowledge, yet much of it is buried in unstructured text and inconsistent terminology, making systematic retrieval and reasoning difficult. Existing tools typically extract narrow, study-specific facts in isolation, failing to preserve the cross-study context required to answer broader scientific questions. Retrieval-augmented generation (RAG) offers a promising way to overcome this limitation by combining large language models (LLMs) with external retrieval, but its effectiveness depends strongly on how domain knowledge is represented. In this work, we develop two retrieval pipelines: a dense semantic vector-based approach (VectorRAG) and a graph-based approach (GraphRAG). Using over 1,000 polyhydroxyalkanoate (PHA) papers, we construct context-preserving paragraph embeddings and a canonicalized structured knowledge graph supporting entity disambiguation and multi-hop reasoning. We evaluate these pipelines through standard retrieval metrics, comparisons with general state-of-the-art systems such as GPT and Gemini, and qualitative validation by a domain chemist. The results show that GraphRAG achieves higher precision and interpretability, while VectorRAG provides broader recall, highlighting complementary trade-offs. Expert validation further confirms that the tailored pipelines, particularly GraphRAG, produce well-grounded, citation-reliable responses with strong domain relevance. By grounding every statement in evidence, these systems enable researchers to navigate the literature, compare findings across studies, and uncover patterns that are difficult to extract manually. More broadly, this work establishes a practical framework for building materials science assistants using curated corpora and retrieval design, reducing reliance on proprietary models while enabling trustworthy literature analysis at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。