arXiv:2601.06117cs.LG2026-01

让AI像物理学家一样自主发现新定律,突破传统模型的数值极限。

The Active Discoverer Framework: Towards Autonomous Physics Reasoning through Neuro-Symbolic LaTeX Synthesis

  • 用最低有效位编码实现无精度损失的数学推理
  • 在10^50量级上仍保持准确,远超传统方法的10^16极限
  • 适合追求高可信科学发现的理论物理研究者

现代人工智能擅长在已知数据分布内进行统计插值,但在理论物理与数学所需的精确推理上根本失败。我们识别出‘浮点墙’——当规模超过10^16时,标准浮点表示和词元化(BPE)导致神经网络外推能力崩溃。为此,我们提出主动发现者框架,一种面向不变性发现的原生数字神经符号架构。核心是NumberNet,一种双胞胎算术变换器,通过最低有效位(LSB)序列编码实现0%精度损失,并支持高达10^50的宇宙尺度外推。为确保物理真实性,引入基于哈密顿的能量下降与对称性分组层,使模型天然遵循诺特定理。主要创新在于符号LaTeX瓶颈:模型通过自回归LaTeX解码器主动推测未知物理变量,将数值‘幻觉’与结构合法数学表达相校验,确保所发现物理规律简洁且人类可读。我们在300亿规模基准和包含50个‘混沌模式’系统扰动的通用物理圣殿上评估,结果表明传统GBDT和基于LLM的架构在宇宙尺度下失效,而主动发现者能高保真自主推导万有引力常数(G)等普适常数。该框架为零幻觉人工智能与真正自主科研代理开辟道路。

原文摘要 · Abstract (English)

Modern artificial intelligence excels at statistical interpolation within seen manifolds but fundamentally fails at the exact reasoning required for theoretical physics and mathematics. We identify the "Float Wall" -- a catastrophic collapse of neural extrapolation at scales beyond $10^{16}$ -- caused by standard floating-point representation and linguistic tokenization (BPE). To resolve this, we introduce the Active Discoverer Framework, a digit-native neuro-symbolic architecture designed for invariant discovery. At its core is NumberNet, a Siamese Arithmetic Transformer that utilizes least-significant-bit (LSB) sequence encoding to achieve 0% precision loss and cosmic-scale extrapolation up to $10^{50}$. To enforce physical grounding, we implement a Hamiltonian-based energy descent and Symmetry Grouping layer, ensuring the model respects Noether's theorem natively. The primary innovation is the Symbolic LaTeX Bottleneck: an active discovery loop where the model is forced to hypothesize unknown physical variables through an autoregressive LaTeX decoder. By reconciling numeric "hallucinations" with structurally valid mathematical expressions, the framework ensures that any discovered physics is parsimonious and human-interpretable. We evaluate this system against a 30-billion scale benchmark and the Universal Physics Pantheon, featuring 50 "Chaos Mode" systemic perturbations. Our results demonstrate that while traditional GBDT and LLM-based architectures collapse at cosmic scales, the Active Discoverer autonomously deduces universal constants such as the Gravitational Constant ($G$) with high fidelity. This framework establishes a path toward zero-hallucination artificial intelligence and truly autonomous scientific research agents.

物理推理神经符号数学发现代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。