提出LoRA-Curve方法,发现并连接了低损耗路径,提升模型不确定性估计。
On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference
- 用分段贝塞尔曲线参数化LoRA空间,构建连续低损路径。
- 在Qwen2.5 7B上验证多模式间存在低损山谷,线性插值有损耗壁垒。
- 适合关注模型不确定性和函数多样性的研究者使用。
尽管低秩适配(LoRA)已成为大语言模型参数高效微调的标配,但对认知不确定性(epistemic uncertainty)的合理估计仍具挑战性。近期研究表明,在LoRA范式中,深度集成等离散多模态方法对泛化提升有限,与深度学习中集成独立最优解通常能改善性能的普遍观察相悖。若能在参数空间中通过连续低损路径连接多个最优解,可进一步增强贝叶斯模型平均(BMA)效果。然而,目前尚不清楚这种结构是否存在,以及是否能揭示局部或离散方法遗漏的功能多样性。本文提出LoRA-Curve,一种在LoRA空间中的分段贝塞尔曲线参数化方法,包含自由配置(联合优化所有控制点)和锚定配置(连接独立微调的最优解)。我们证明了路径上的损失具有路径连续性和Lipschitz正则性。在推理与分类任务中,基于Qwen2.5 7B的实证表明,线性插值会遭遇损失壁垒,而我们的锚定多段曲线可实现跨独立最优解的连续低损连接。结合平坦极小值扰动与詹森-香农散度正则化,LoRA-Curve显著提升了预测分布的互信息,同时保持性能不下降,将连续参数空间遍历与功能多样性联系起来。
原文摘要 · Abstract (English)
While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic uncertainty remains challenging. Recent results in the LoRA regime suggest that discrete multi-mode approaches such as deep ensembles offer little benefit over single-mode methods. This contradicts broader observations in deep learning, where ensembling independent optima typically improves generalization, and linking these modes through continuous low-loss valleys further enhances Bayesian model averaging (BMA). Whether such structure exists in the LoRA space and whether it yields functional diversity missed by local or discrete methods has not been studied. We introduce LoRA-Curve, a segmented Bézier curve parameterization in the LoRA space, with two variants: a free configuration that jointly optimizes all control points, and an anchored configuration that connects independently fine-tuned LoRA optima. We prove pathwise continuity and Lipschitz regularity of the loss along the curve and empirically show, across reasoning and classification benchmarks with Qwen2.5 7B, that linear interpolation encounters loss barriers, while our anchored multi-segment curves connect independent optima through continuous low-loss valleys. Combined with flat-minima perturbations and a Jensen-Shannon divergence regularizer, LoRA-Curve yields measurably higher mutual information of the predictive distribution without sacrificing performance, and links continuous parameter-space traversal to functional diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。