arXiv:2506.21763cs.AI2025-06被引 1

用历史演化树结构提升科学命题的验证与推理能力

THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?

  • 构建领域专属科学演进树,通过思维-引用-验证闭环确保每步有据可依
  • 在图补全任务中相比传统引文网络提升8%~14%命中率,预测未来进展提升近10%
  • 适用于科学发现评估、论文重要性判断,尤其适合需要深度溯源的研究者

大型语言模型正在加速科学构想生成,但对海量、常表层的AI生成命题进行新颖性与事实准确性的严格评估是关键瓶颈;人工验证效率过低。现有验证方法不足:独立使用大模型作为验证器易产生幻觉且缺乏领域知识(我们发现其在特定领域对相关文献的知晓率仅40%),而传统引文网络缺乏显式因果关系,叙述性综述又无结构。这凸显核心挑战:缺乏结构化、可验证且具因果关联的科学演化历史数据。为此,我们提出 extbf{THE-Tree}(Technology History Evolution Tree),一个从科学文献中构建领域专属演化树的计算框架。THE-Tree采用搜索算法探索演化路径,在节点扩展时运用创新的“思考-表述-引用-验证”流程:大模型提出潜在进展并引用文献;关键在于,每个演化链路均通过恢复的自然语言推理机制,对所引文献进行质询,以检验逻辑连贯性与证据支持,确保每一步均有据可依。我们在多个领域构建并验证了88个THE-Trees,发布基准数据集,包含最多71,000次事实验证,覆盖27,000篇论文,以促进后续研究。实验表明:(i) 在图补全任务中,THE-Tree相较传统引文网络在多个模型上将命中率(hit@1)提升8%至14%;(ii) 预测未来科学发展时,命中率提升近10%;(iii) 与其他方法结合后,评估重要科学论文的性能提升接近100%。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are accelerating scientific idea generation, but rigorously evaluating these numerous, often superficial, AI-generated propositions for novelty and factual accuracy is a critical bottleneck; manual verification is too slow. Existing validation methods are inadequate: LLMs as standalone verifiers may hallucinate and lack domain knowledge (our findings show 60% unawareness of relevant papers in specific domains), while traditional citation networks lack explicit causality and narrative surveys are unstructured. This underscores a core challenge: the absence of structured, verifiable, and causally-linked historical data of scientific evolution.To address this,we introduce \textbf{THE-Tree} (\textbf{T}echnology \textbf{H}istory \textbf{E}volution Tree), a computational framework that constructs such domain-specific evolution trees from scientific literature. THE-Tree employs a search algorithm to explore evolutionary paths. During its node expansion, it utilizes a novel "Think-Verbalize-Cite-Verify" process: an LLM proposes potential advancements and cites supporting literature. Critically, each proposed evolutionary link is then validated for logical coherence and evidential support by a recovered natural language inference mechanism that interrogates the cited literature, ensuring that each step is grounded. We construct and validate 88 THE-Trees across diverse domains and release a benchmark dataset including up to 71k fact verifications covering 27k papers to foster further research. Experiments demonstrate that i) in graph completion, our THE-Tree improves hit@1 by 8% to 14% across multiple models compared to traditional citation networks; ii) for predicting future scientific developments, it improves hit@1 metric by nearly 10%; and iii) when combined with other methods, it boosts the performance of evaluating important scientific papers by almost 100%.

科学推理知识图谱验证机制演化树

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。