arXiv:2606.07908cs.LG2026-06

通过分层导数控制提升模型在不同数据量下的稳定性和泛化能力

Layer-wise Derivative Controlled Networks Achieve Competitive Accuracy and Gradient Stability Across Data Regimes

论文配图:Layer-wise Derivative Controlled Networks Achieve Competitive Accuracy and Gradient Stability Across Data Regimes
图 1 · 摘自论文原文
  • 采用分层导数惩罚机制,动态调节每层梯度以抑制噪声
  • 在5%至100%数据下均保持优于基线的准确率,梯度尾比稳定在1.01–1.02
  • 适用于小样本、低质量表示场景,适合追求鲁棒性的研究者

基于链式法则(CR)的导数控制网络结合了三次多项式层与轻量级前向模式逐层雅可比惩罚(DREG)。本文评估了CR在不同数据条件下的泛化性能。通过消融实验发现,DREG系数调度的最优退火范围取决于表示噪声水平。在Pima糖尿病数据集上,CR在低数据量下表现优异,并在5%至100%训练数据范围内持续保持显著准确率优势,其梯度尾比稳定在~1.01–1.02(而ReLU网络为1.07–1.09)。在SST-5上的扩展实验显示,在冻结嵌入与BERT微调场景中均达到或超越现有基线,且使用更少训练数据即实现领先结果。统计检验表明,CR在两个数据集上均显著优于最强公开基线(p < 0.05)。这些结果证明,分层导数控制引入了低频、稳定的表征结构归纳偏置,可在表格与NLP领域、不同数据量及表示质量下稳健泛化。梯度尾比成为可靠的无标签泛化能力诊断指标。

原文摘要 · Abstract (English)

Derivative-controlled networks based on ChainzRule (CR) combine cubic polynomial layers with a lightweight forward-mode per-layer Jacobian penalty (DREG). In this second paper of a multi-part series, we evaluate the generalization properties of CR across data regimes. We ablate the shape of the DREG coefficient schedule, demonstrating that the optimal annealing range depends on representation noise. On the Pima Diabetes dataset, CR achieves strong low-data performance and maintains a consistent accuracy advantage over baselines from 5\% to 100\% training data, supported by exceptionally stable gradient tail ratios ($\sim$1.01--1.02 vs. 1.07--1.09 for ReLU networks). Extensions to SST-5 show competitive or superior results in both frozen-embedding and BERT fine-tuned regimes, including outperforming prior BERT baselines despite substantially less training data. These results are statistically significant: CR achieves superior accuracy over the strongest published baselines we could identify on both datasets ($p < 0.05$). These results establish that layer-wise derivative control induces a structural inductive bias toward low-frequency, stable representations that generalizes robustly across tabular and NLP domains, data volumes, and representation qualities. The gradient tail ratio serves as a reliable, label-free diagnostic of generalization capability.

导数控制梯度稳定性小样本学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。