通过QR分解挖掘视觉模型潜空间,实现更优的领域泛化分割。
RecycleLoRA: Rank-Revealing QR-Based Dual-LoRA Subspace Adaptation for Domain Generalized Semantic Segmentation
- 用秩揭示QR分解挖掘预训练模型的潜在子空间结构。
- 主适配器独立学习多样特征,性能接近完整模型。
- 双适配器分工协同,无需额外正则化,推理无延迟。
领域泛化语义分割(DGSS)旨在跨未见目标域保持鲁棒性能。视觉基础模型(VFMs)蕴含丰富的多域知识,可提升泛化能力。然而,现有方法对VFMs中丰富子空间结构的主动利用仍不足,多数聚焦于保留预训练知识,且LoRA组件常面临表征多样性低与参数利用率差的问题。本文提出RecycleLoRA,通过秩揭示QR分解(RRQR)系统性挖掘VFMs的子空间结构,增强LoRA的表征能力。主适配器利用RRQR识别的微小子空间方向学习多样化、独立特征,单独使用即达竞争性表现;同时引入子适配器,对主要方向进行轻量级微调,与主适配器形成互补提升。该设计使双适配器学习不同表示,无需额外正则化损失。基于RRQR的预训练子空间结构系统性利用,显著提升领域泛化性能。RecycleLoRA在合成到真实及真实到真实泛化任务上均达到当前最优表现,无需复杂架构或增加推理延迟。
原文摘要 · Abstract (English)
Domain Generalized Semantic Segmentation (DGSS) aims to maintain robust performance across unseen target domains. Vision Foundation Models (VFMs) offer rich multi-domain knowledge that can enhance generalization. However, strategies for actively exploiting the rich subspace structures within VFMs remain under-explored, with many existing methods focusing primarily on preserving pre-trained knowledge. Furthermore, their LoRA components often suffer from limited representational diversity and inefficient parameter utilization. We propose RecycleLoRA, which addresses both challenges by employing Rank-Revealing QR Decomposition (RRQR) to systematically exploit VFM's subspace structures and enhance LoRA's representational richness. Our main adapter leverages minor subspace directions identified by RRQR to learn diverse and independent features, achieving competitive performance even when used alone. We further introduce a sub adapter that carefully refines major directions with minimal adjustments, providing complementary improvements to the main adapter's strong baseline performance. This design enables the dual adapters to learn distinct representations without requiring additional regularization losses. Our systematic exploitation of pre-trained subspace structures through RRQR-based initialization leads to superior domain generalization performance. RecycleLoRA achieves state-of-the-art performance on both synthetic-to-real generalization and real-to-real generalization tasks without complex architectures or additional inference latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。