arXiv:2607.28684cs.AI2026-07

检验语言模型先验在科学公式发现中的真实作用。

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

论文配图:Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
图 1 · 摘自论文原文
  • 用固定词汇库构建无语义基线,对比语言模型生成候选
  • 多数任务已由固定词汇覆盖,模型先验贡献有限
  • 仅在词汇覆盖受扰时,先验才显著提升表现

现有科学公式发现基准多基于公开可得的经典方程,难以区分模型是数据中发现规律还是单纯记忆训练集。LSR-Synth通过引入新型合成项并筛选任务的新颖性、可解性和科学合理性来缓解此问题。本文聚焦更窄的测量问题:语言模型提供的科学先验能否区分于不依赖任务语义的常规算子搜索?我们构建了一个基于公开来源固定词汇的无语义基线,通过语义屏蔽、知识库弱化和匹配算子族剔除评估候选覆盖度。在当前任务快照、搜索预算与评分协议下,固定词汇已覆盖大部分任务,而语言模型生成候选极少扩展可解实例;其边际贡献仅在词汇覆盖被选择性破坏时显现。严格分布外评估虽降低所有方法成功率,但不改变这一关系。研究不否定LSR-Synth对完整公式记忆的防控能力,也不意味着语言模型先验普遍无效,而是支持更有限结论:当前多数任务仍适合评估未见表达式的拟合与重组,但不足以独立识别超出固定搜索空间的先验贡献。

原文摘要 · Abstract (English)

Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.

公式发现语言模型先验知识评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。