arXiv:2608.04393cs.LG2026-08

提出保留公式结构的残差学习方法,提升土壤侵蚀预测对因子退化的鲁棒性。

When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision

论文配图:When Proxy Prediction Becomes Equation Reconstruction: Diagnostics and Residual Learning for Factor-Derived Proxy Supervision
图 1 · 摘自论文原文
  • 用残差框架保持公式估计作为预测锚点,学习上下文修正项
  • 在因子退化下,相比直接预测,误差降低32%以上,尾部误差下降显著
  • 适合需要高鲁棒性的科学建模场景,如环境预测与数据受限领域

科学机器学习常依赖已知领域因子生成的代理目标,当这些因子同时作为模型输入时,高预测准确率可能反映的是对生成代理方程的重建,而非对退化因子信息的鲁棒性。本文在土壤可蚀性因子 $K$ 受控退化条件下,研究了基于 RUSLE 的土壤损失代理预测问题。提出诊断框架,结合退化公式参考、经典树模型基线、匹配的直接与公式特征预测器、上下文消融分析、尾部误差分析及退化鲁棒性评分。进一步提出 RASPL,一种保留公式的残差学习框架:以退化公式估计为预测锚点,学习自适应门控的上下文修正。RASPL 显著优于匹配的直接预测,在退化和尾部误差方面均表现更优。其中,紧凑统计编码器实现最高宏平均 $R^2$ 与最低计算成本,卷积编码器则具备最强退化鲁棒性与最低 Tail95 MAE。结果确立公式保留为因子衍生代理监督下的核心设计原则。

原文摘要 · Abstract (English)

Scientific machine learning often relies on proxy targets computed from known domain factors when direct observations are limited. When those same factors are used as model inputs, however, high predictive accuracy may reflect reconstruction of the proxy-generating equation rather than robustness to degraded factor information. We study this problem in RUSLE-derived soil-loss proxy prediction under controlled degradation of the soil-erodibility factor $K$. We introduce a diagnostic framework that combines degraded-formula references, classical tree-based baselines, matched direct and formula-feature predictors, contextual ablations, tail-error analysis, and degradation robustness scoring. We then propose RASPL, a formula-preserving residual framework that retains the degraded formula estimate as the prediction anchor and learns an adaptively gated contextual correction. RASPL substantially outperforms matched direct prediction and provides stronger degradation and tail robustness than treating the formula estimate as an ordinary input feature. Within RASPL, a compact statistical encoder achieves the highest macro-averaged $R^2$ and lowest computational cost, whereas a convolutional encoder achieves the strongest degradation robustness and lowest Tail95 mean absolute error (MAE). These results establish formula preservation as the central design principle for robust learning from factor-derived proxy targets.

科学机器学习代理监督鲁棒性残差学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。