提出新方法评估大模型适配器的实际有效秩,发现其常远低于设计值。
Spectral Rank Certification for Foundation Model Adapters

- 基于谱统计推断框架,通过贝叶斯先验与卡方散度计算有限样本下有效秩
- 实测26个公开适配器共684模块,有效秩普遍远低于名义秩,与能量保留率不同
- 适用于大模型微调审计、适配器设计优化,尤其关注秩压缩与统计可信度
名义LoRA秩是设计参数;经校准的谱证据则是独立的推断量。本文构建了针对公共基础模型适配器有效秩结构的有限样本推断框架。理论核心为固定维高斯秩一参考实验的精确卡方散度,对未知信号方向在旋转不变先验下积分。由此导出可计算的有限样本Le Cam界、数值截断的显式余项界,以及矩形版Baik-Ben Arous-Péché (BBP) 极限。紧凑流形拉普拉斯展开表明,有限样本似然证据还依赖于主谱隙,因子形式为 $s_1^{|m-n|} imes\prod_{i\ge2}(s_1^2-s_i^2)$,推动聚类奇异值的联合校准。基于此,我们提出一个用于PEFT LoRA适配器的实证-零假设工作流:因子重构、蒙特卡洛p值、分阶段与块测试、模块级与语料级BH报告。对26个公开适配器、684个模块、六种架构族、31,770条公共检查点谱数据的审计显示,校准后有效秩通常远小于名义秩,且系统性偏离95%能量保留。以$ n=24 $样本的RoBERTa-RTE片段为例,展示了从校准秩到任务评估的完整路径,未将其视为效用研究。主要实证发现为:校准有效秩通常显著低于名义秩,而能量保留与统计意外回答的是不同问题。
原文摘要 · Abstract (English)
Nominal LoRA rank is a design parameter; calibrated spectral evidence is a separate inferential quantity. This article develops a finite-sample framework for inferring effective rank structure in public foundation-model adapters. The theoretical core is an exact chi-square divergence for the fixed-dimensional Gaussian rank-one reference experiment, with an unknown signal direction integrated under a rotation-invariant reference prior. The resulting series yields a computable finite-sample Le Cam bound at concrete layer sizes, an explicit remainder bound for numerical truncation, and the rectangular Baik-Ben Arous-Peche (BBP) limit. A compact-manifold Laplace expansion shows that finite-sample likelihood evidence also depends on leading spectral gaps through the factor $s_1^{|m-n|}\prod_{i\ge2}(s_1^2-s_i^2)$, motivating joint calibration of clustered singular values. Building on these results, we introduce an empirical-null workflow for PEFT LoRA adapters: factor reconstruction, Monte Carlo $p$-values, stagewise and block testing, and module-wise and corpus-level BH reporting. In an audit of 26 public adapters, 684 modules, six architecture families, and 31,770 public-checkpoint spectra rows, calibrated effective rank is typically much smaller than nominal rank and differs systematically from 95\% energy retention. A measured RoBERTa-RTE slice on $n=24$ examples illustrates the measurement path from calibrated ranks to task evaluation, without treating the slice as a utility study. The main empirical finding is that calibrated effective rank is usually far below nominal rank, and that energy retention and statistical surprise answer different questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。