研究发现:小规模科学知识图谱可替代完整图生成有效假设。
The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

- 通过压缩知识图谱的密度、结构和语义层级,测试其对假说生成的影响。
- 仅用前k个关键三元组即可接近全图效果,即使关键结论被隐藏。
- 随机或拓扑结构子图也能保留核心信号,适合高效推理场景。
知识图谱(KG)能为语言模型提供结构化科学背景,但尚不清楚哪些图事实真正影响生成的假说。本研究在Mistral-7B、Llama-3.1-70B和Gemini 2.5 Flash上,针对电池材料的假说生成任务,通过改变局部知识图谱的密度、本体丰富度、拓扑结构和控制结构进行扰动,并采用提供图与固定参考双重指标评估输出。结果显示,知识图谱的作用具有选择性且依赖模型:图上下文会改变输出,但无图输入也可通过模型先验恢复大量图内容。紧凑的top-k子图常能逼近完整图表现,即使声称结果三元组被移除。同时,压缩并非依赖单一语义排序规则,随机或基于拓扑的子集同样可恢复大部分信号。这些结果支持冗余感知的压缩型知识图谱假设:有用的知识信号通常可从小型、科学结构化的子图中恢复,无需依赖完整局部图。
原文摘要 · Abstract (English)
Knowledge graphs (KGs) can provide structured scientific context to language models, but it remains unclear which graph facts actually shape the generated hypotheses. We study KG-guided hypothesis generation for battery materials across Mistral-7B, Llama-3.1-70B, and Gemini 2.5 Flash. We perturb local KGs by varying density, ontology richness, topology, and control structure, and evaluate outputs with both provided-graph and fixed-reference metrics. Across models, KG utility is selective and model-dependent: graph context changes outputs, but no-KG outputs also recover substantial graph content from model priors. Compact top-k subgraphs often approximate full-KG behavior, including when claimed-outcome triples are held out. At the same time, compression is not unique to one semantic ranking rule, random and topology-based subsets can also recover much of the signal. These results support a redundancy-aware Compressive KG hypothesis: useful KG signal is often recoverable from compact, scientifically structured subgraphs rather than requiring the full local graph.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。