高维微调中,少数关键方向主导优化,解释小种群也能高效训练。
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
- 通过曲率分析发现优化方向高度集中于少数维度。
- 小种群(约30)即可有效微调百亿参数模型,且奖励先升后降。
- 适用于理解进化策略与随机优化在大模型中的有效性。
权重扰动进化策略(ES)可在极小种群(如 $Nacksim30$)下微调百亿参数语言模型,违背经典零阶维度诅咒直觉。我们还观察到另一现象:固定超参数下,随机微调奖励常先上升、达峰后下降,无论在ES还是GRPO中均成立。我们认为二者均源于微调景观的共同几何特性:曲率低维性。少数高曲率维度主导优化,导致(i)不同时间尺度下的先升后降动态,由最小二次随机上升模型捕捉;(ii)退化的优化更新,即多种随机扰动共享相同改进方向。以ES为几何探测工具,在GSM8K、ARC-C和WinoGrande数据集上对Qwen2.5-Instruct(0.5B–7B)进行分析,结果表明奖励提升扰动在不同规模下仍可通过小种群访问。这些发现调和了ES可扩展性与非单调训练动态的矛盾,提示高维微调可能比最坏情况理论允许的优化方法更广泛。
原文摘要 · Abstract (English)
Weight-perturbation evolution strategies (ES) can fine-tune billion-parameter language models with surprisingly small populations (e.g., $N\!\approx\!30$), contradicting classical zeroth-order curse-of-dimensionality intuition. We also observe a second seemingly separate phenomenon: under fixed hyperparameters, the stochastic fine-tuning reward often rises, peaks, and then degrades in both ES and GRPO. We argue that both effects reflect a shared geometric property of fine-tuning landscapes: they are low-dimensional in curvature. A small set of high-curvature dimensions dominates improvement, producing (i) heterogeneous time scales that yield rise-then-decay under fixed stochasticity, as captured by a minimal quadratic stochastic-ascent model, and (ii) degenerate improving updates, where many random perturbations share similar components along these directions. Using ES as a geometric probe on fine-tuning reward landscapes of GSM8K, ARC-C, and WinoGrande across Qwen2.5-Instruct models (0.5B--7B), we show that reward-improving perturbations remain empirically accessible with small populations across scales. Together, these results reconcile ES scalability with non-monotonic training dynamics and suggest that high-dimensional fine-tuning may admit a broader class of viable optimization methods than worst-case theory implies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。