提出共享符号主干方法,让多输出回归模型保持物理一致性。
Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression

- 用神经演化法搜索共享的符号基底,通过稀疏加乘读出适配多输出。
- 在弱可识别物理因子场景下,显著提升多输出一致性,独立回归无法做到。
- 适合需要提取共享机制、结构稀疏的物理系统建模任务。
符号回归能生成解析表达式,但通常逐输出单独应用,这在过程系统中受限,因状态变量常通过共享物理参数耦合。独立符号回归虽可得高精度个体方程,但难以整合为统一模型。本文提出一种神经演化符号回归方法,用于耦合多输出系统。该方法搜索共享符号主干:一组通过稀疏加性或乘性读出被多个输出复用的潜在符号单元。离散模型结构通过变异和交叉演化,连续参数则通过梯度下降优化并传递给后代。在具有已知真值的基准数据集及生物油水热液化产率案例上评估。结果表明,耦合并非普遍降低预测误差;其主要价值在于强制并诊断跨输出一致性,当物理共享因子嵌入隐式表达且数据中弱可识别时(如朗缪尔-欣谢尔伍德和位点覆盖分母),独立PySR无法弥合一致性差距或恢复相同共享形式。相反,在每个输出已可识别的情形(如Van de Vusse基准),独立回归表现相当或更优。所提框架非通用预测器,而是结构化的共享机制提取工具,其价值在目标结构稀疏、共享、弱可识别或受闭合约束时最高。
原文摘要 · Abstract (English)
Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, where state variables are often coupled through shared physical parameters. Independent symbolic regression can give accurate individual equations that are difficult to interpret as one model. We present a neuro-evolutionary symbolic regression method for coupled multi-output systems. The method searches for a shared symbolic backbone: a set of latent symbolic units that is discovered once and reused by several outputs through sparse additive or multiplicative read-outs. The discrete model structure is evolved by mutation and crossover, whereas the continuous parameters are tuned by gradient descent and inherited by the offspring. The method is assessed on a set of benchmarks with known ground truth and on a hydrothermal liquefaction yield case. The results show that coupling is not a general route to lower prediction error. Its main contribution is the enforcement and diagnosis of cross-output consistency when a physically shared factor is embedded in a latent expression and is weakly identifiable from the data. This occurs for Langmuir-Hinshelwood and site-coverage denominators, for which independent PySR does not close the consistency gap or recover the same shared form. Conversely, when each output is already identifiable, as in the Van de Vusse benchmark, independent symbolic regression matches or improves the coupled model. The proposed framework, rather than a general purpose predictor, is a structured shared-mechanism extractor. Its value is highest when the target structure is sparse, shared, weakly identifiable or constrained by closure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。