arXiv:2605.30333cs.CL2026-05被引 1

用引文和定理结构联合生成可信的未来数学命题。

COMPOSE: Composing Future Theorems from Citations and Formal Structure

  • 结合引文网络与定理依赖图,双重约束生成过程
  • 在47,000个未来论文中检索表现超越基线
  • 适合数学研究者和形式化推理系统开发者

一个可信的未来数学命题需满足两个条件:延续已有研究方向,并遵守形式逻辑的推导约束。现有方法通常仅考虑其一,导致结论或缺乏依据或动机不足。本文提出基于科学引文图与形式定理依赖图的联合生成框架COMPOSE,通过双图结构引导语言模型生成合理的未来定理。我们构建了10.8万对科学-形式图数据集(来自arXiv与Mathlib),并设立2024–2025年47,000个未来论文的基准测试。实验表明,COMPOSE在真实未来论文检索任务中优于强基线,在大模型评估中整体表现最佳,生成内容更扎实且数学丰富。结果证明,结合科学语境与形式结构能显著提升未来数学命题生成质量。

原文摘要 · Abstract (English)

A plausible future mathematical claim must satisfy two constraints: it should follow the direction of prior work and respect the formal dependencies that constrain what can validly follow. Existing approaches typically model only one of these sources, producing claims that are either weakly grounded or insufficiently motivated. We introduce grounded future mathematical generation, where the goal is to generate a plausible future theorem-like claim for an anchor paper using two complementary sources of context: its scientific citation graph and aligned formal theorem dependency graph. To address this setting, we propose COMPOSE, a dual-graph framework that conditions a language model on both scientific citation context and formal theorem structure. To support this setting, we construct a dataset of 108K paired scientific-formal graph examples from arXiv and Mathlib, together with a benchmark of 47K future papers from 2024--2025. Experiments show that COMPOSE outperforms strong baselines on retrieval to real future papers and achieves the best overall performance under LLM-judge evaluation, producing more grounded and mathematically richer outputs. These results show that future mathematical generation benefits from combining scientific context with formal structure. Project page is available at https://david-busbib.github.io/COMPOSE-page/.

数学生成定理推理双图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。