arXiv:2506.08267cs.LGcs.AI2025-06

用可解释的神经网络找简洁数学公式,比现有方法更准更省

Sparse Interpretable Deep Learning with LIES Networks for Symbolic Regression

  • 设计固定结构的LIES网络,用对数、指数等可解释激活函数建模
  • 通过特殊采样和损失函数训练,使结果既准确又稀疏(仅含必要项)
  • 适合需要简洁可读公式的科研与工程场景

符号回归(SR)旨在发现能精确描述数据的闭式数学表达式,提供黑箱模型所不具备的可解释性与分析洞见。现有方法多依赖种群搜索或自回归建模,面临可扩展性差和符号不一致的问题。我们提出LIES(对数、恒等、指数、正弦)架构,一种具有可解释基础激活函数的固定神经网络,用于建模符号表达式。通过采用合适的过采样策略和定制化损失函数训练,促进稀疏性并防止梯度不稳定。训练后进一步应用剪枝策略,将学习到的表达式简化为紧凑形式。在多个符号回归基准测试中,LIES框架持续生成稀疏且精确的符号公式,优于所有基线方法。消融实验验证了各设计组件的重要性。

原文摘要 · Abstract (English)

Symbolic regression (SR) aims to discover closed-form mathematical expressions that accurately describe data, offering interpretability and analytical insight beyond standard black-box models. Existing SR methods often rely on population-based search or autoregressive modeling, which struggle with scalability and symbolic consistency. We introduce LIES (Logarithm, Identity, Exponential, Sine), a fixed neural network architecture with interpretable primitive activations that are optimized to model symbolic expressions. We develop a framework to extract compact formulae from LIES networks by training with an appropriate oversampling strategy and a tailored loss function to promote sparsity and to prevent gradient instability. After training, it applies additional pruning strategies to further simplify the learned expressions into compact formulae. Our experiments on SR benchmarks show that the LIES framework consistently produces sparse and accurate symbolic formulae outperforming all baselines. We also demonstrate the importance of each design component through ablation studies.

符号回归可解释模型稀疏学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。