低秩适配让视觉语言对齐更稳定高效,还能减少幻觉。
Dive Into the Implicit Biases of Low-rank Vision-language Alignment

- 用低秩方式调整大模型,降低计算开销。
- 在多数基准上表现优于全参数更新,且减少幻觉。
- 适合关注模型稳定性与效率的研究者。
视觉语言对齐通常被视为需全参数微调的预训练阶段。本文挑战这一观点,研究在该阶段对大语言模型采用低秩适配的效果。发现低秩对齐不仅降低计算成本,还在多数基准上超越全参数对齐。通过系统分析,发现其引入隐式偏差:使模型行为从易幻觉转为保守,并保持视觉特征的逐标记线性可分性(称作LS灾变)。几何上,低秩对齐模型具有更均一、结构稳定的视觉表征,保留模态特异性知识而非过早融合实体语义。理论上,证明了低秩对齐倾向于选择梯度平坦的参数子空间和对抗扰动的特征子空间,为结构保持提供理论解释。实验涵盖超过100种对齐配置、三类低秩算子及多种秩、编码器等设置。
原文摘要 · Abstract (English)
Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates. We challenge this view and investigate what happens when low-rank adaptation is applied to the LLM during this stage instead. We find that low-rank alignment not only reduces computational costs but also outperforms full-parameter alignment on most benchmarks. To understand this phenomenon, we systematically characterize the implicit biases introduced by low-rank adaptation during alignment. Empirically, we find that low-rank alignment shifts model behavior from hallucinatory to conservative and preserves per-token linear separability of visual features that full-parameter alignment disrupts, a phenomenon we term LS-curse. Geometrically, low rank aligned models exhibit more homogeneous and structurally stable visual representations, maintaining modality-specific knowledge rather than prematurely fusing entity-level semantics. Theoretically, we establish two theorems showing that low-rank alignment induces preferences for parameter subspaces with flat gradients and feature subspaces robust to perturbations, providing a principled explanation for the observed structure-preserving behavior. Extensive experiments cover ablation over 100 alignment configurations, three families of low-rank operators, and various rank, encoder, and other settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。