用低秩正交变换提升模型微调的几何一致性与效率
LoCO: Low-rank Compositional Rotation Fine-tuning

- 通过低秩反对称矩阵构造正交旋转链,保持预训练表示结构
- 支持全并行计算,高维特征空间下仍保持低复杂度与小误差
- 适用于视觉、语言、扩散模型等多领域,性能优于或媲美主流方法
参数高效微调(PEFT)已成为大规模基础模型在自然语言处理和计算机视觉中适配的关键技术。现有方法如低秩适应虽通过低秩权重更新实现参数效率,但难以保持预训练表征的几何结构。我们提出一种新型PEFT方法——低秩组合正交微调(LoCO),通过低秩反对称矩阵构建正交变换,并采用组合旋转链。我们设计了一种近似方案,实现组合旋转的完全并行计算,使该方法在高维特征空间中具备实用性。该方法在保持低计算复杂度的同时,维持正交性且逼近误差可控。我们在扩散变压器微调、视觉变压器适配及语言模型微调等多个领域验证了其有效性,结果表明,该方法在性能上优于或媲美现有正交与非正交方法。
原文摘要 · Abstract (English)
Parameter-efficient fine-tuning (PEFT) has emerged as an critical technique for adapting large-scale foundation models across natural language processing and computer vision. While existing methods such as low-rank adaptations achieve parameter efficiency via low-rank weight updates, they are limited in their ability to preserve the geometric structure of pretrained representations. We introduce Low-rank Compositional Orthogonal fine-tuning (LoCO), a novel PEFT method that constructs orthogonal transformations through low-rank skew-symmetric matrices and compositional rotation chains. We propose an approximation scheme that enables fully parallel computation of compositional rotations, making the approach practical for high-dimensional feature spaces. Our method maintains low computational complexity while maintaining orthogonality with controlled approximation error. We validate LoCO across diverse domains, including diffusion transformer fine-tuning, vision transformer adaptation, and language model adaptation. Our method demonstrates superior or competitive performance compared to both existing orthogonal and non-orthogonal methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。