提出新指标InTrain,用理论方法评估模型是否易训练。
InTrain: Intrinsic Trainability for Zero-Cost Neural Architecture Search
- 从几何容量和优化韧性两方面量化模型可训练性
- 在多个基准上表现媲美顶尖集成方法
- 适合快速筛选高潜力网络结构的工程师
零成本神经架构搜索有望在无需昂贵训练的情况下高效发现高性能网络。然而,现有零成本代理依赖碎片化启发式方法,未能捕捉核心问题:是什么让一个架构具备可训练性?本文提出内在可训练性(InTrain),一种统一的理论代理,将可训练性形式化为由几何容量与优化韧性两个协同组件共同产生的架构不变量。通过分析神经信息处理过程来实现内在可训练性:几何容量通过激活协方差特征谱的参与度比率量化,反映表示流形的有效维度;优化韧性则通过累积梯度健康度衡量,评估反向传播在深层网络中的鲁棒性。InTrain通过尺度无关的乘法耦合整合这两个维度,我们假设这种非加性关系对捕捉协同效应至关重要。在标准神经架构搜索基准和搜索空间上的大量实验表明,InTrain的排名相关性达到与最先进集成代理相当的水平,并优于其他单指标方法。
原文摘要 · Abstract (English)
Training-free neural architecture search promises efficient discovery of high-performance networks without costly training. However, existing zero-cost proxies rely on fragmented heuristics that fail to capture the fundamental question: what makes an architecture trainable? This paper introduces Intrinsic Trainability (InTrain), a unified theoretical proxy that formalizes trainability as an architectural invariant emerging from two synergistic components: geometric capacity and optimization resilience. We operationalize intrinsic trainability through analysis of neural information processing. Geometric capacity is quantified via the participation ratio of activation covariance eigenspectrum, capturing the effective dimensionality of representation manifolds. Optimization resilience is measured through cumulative gradient health, assessing the robustness of backpropagation across network depth. InTrain synthesizes these dimensions through a scale-invariant multiplicative coupling, which we hypothesize is essential for capturing their synergistic, non-additive relationship. Extensive experiments on standard NAS benchmarks and search spaces demonstrate that InTrain achieves ranking correlations on par with state-of-the-art ensemble-based proxies and outperforms other single-metric methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。