arXiv:2508.02126cs.LGstat.ML2025-08被引 1

通过结构化设计提升模型训练稳定性与可解释性

Understanding Learning Dynamics Through Structured Representations

  • 引入带约束路径和自适应修正的增强变换层
  • 实现更平稳优化、更强鲁棒性及可扩展深度表现
  • 适合关注模型训练机制与可解释架构设计的研究者

尽管现代深度网络展现出卓越的泛化能力,其训练动态仍缺乏深入理解,常依赖经验调参而非架构洞察。本文探究内部结构选择如何影响学习系统行为。基于已有简单架构约束的工作,我们拓展研究结构对收敛性、泛化性和适应性的广泛影响。提出一类包含约束路径与自适应修正的增强变换层,分析其对梯度流动、谱敏感性和固定点行为的影响,揭示了提升训练稳定性和表征规整性的机制。理论分析结合合成与结构化任务的实证研究,验证了改进的鲁棒性、更平滑的优化过程以及可扩展的深度表现。不推荐固定模板,而是强调可解释设计原则,引导学习行为。研究支持架构设计不仅是性能调优,更是塑造可扩展、可信神经系统学习动态的关键维度。

原文摘要 · Abstract (English)

While modern deep networks have demonstrated remarkable versatility, their training dynamics remain poorly understood--often driven more by empirical tweaks than architectural insight. This paper investigates how internal structural choices shape the behavior of learning systems. Building on prior efforts that introduced simple architectural constraints, we explore the broader implications of structure for convergence, generalization, and adaptation. Our approach centers on a family of enriched transformation layers that incorporate constrained pathways and adaptive corrections. We analyze how these structures influence gradient flow, spectral sensitivity, and fixed-point behavior--uncovering mechanisms that contribute to training stability and representational regularity. Theoretical analysis is paired with empirical studies on synthetic and structured tasks, demonstrating improved robustness, smoother optimization, and scalable depth behavior. Rather than prescribing fixed templates, we emphasize principles of tractable design that can steer learning behavior in interpretable ways. Our findings support a growing view that architectural design is not merely a matter of performance tuning, but a critical axis for shaping learning dynamics in scalable and trustworthy neural systems.

架构设计训练动态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。