arXiv:2608.07749cs.CVcs.AI2026-08

提出LoRSA框架,提升生物医学视觉模型在罕见场景下的适应能力。

LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

  • 分离全局与残差更新路径,动态学习互补的低秩参数调整。
  • 在两个外部数据集上分别提升3.09%和2.15%的F1分数,优于现有方法。
  • 适合资源受限下需跨域泛化的生物医学图像任务开发者使用。

参数高效微调可在有限计算资源下适配视觉基础模型至生物医学任务,但单一低秩更新会将所有任务特异性变化限制于单一狭窄参数子空间,可能阻碍模型同时捕捉全局共享结构与局部残差方向,影响对未见成像域的泛化能力。本文提出LoRSA,一种全局-残差自适应框架,联合学习密集低秩分量与动态结构稀疏低秩分量:前者捕捉全局协调的任务适应,后者提供随训练演化的支持结构的互补残差修正。我们分析了该分解的表征能力、近似性质、秩结构及奇异子空间互补性。在使用DINOv3-Base进行四类乳腺密度分类的任务中,以VinDr-Mammo为源域,MammosighTR和RSNA为未见外部域,LoRSA在内部验证集保持竞争力,并在两个目标数据集上取得最优外部宏F1值,较最强对比方法分别提升2.15个百分点(MammosighTR)和3.09个百分点(RSNA)。权重矩阵分析显示,每一分量约92%的能量位于对方双边奇异子空间之外,表明两分量学习到的更新方向高度互补。结果表明,将适应能力划分为全局与残差路径可显著提升参数高效微调后生物医学视觉模型的跨域泛化性能。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. This restriction may prevent the model from simultaneously representing globally shared task structure and localized residual directions required for generalization to unseen imaging domains. We introduce LoRSA, a global--residual adaptation framework that jointly learns a dense low-rank component and a dynamically structured-sparse low-rank component. The dense component captures globally coordinated task adaptation, while the structured component provides complementary residual corrections whose support evolves during training. We characterize the representational capacity, approximation properties, rank structure, and singular-subspace complementarity of this decomposition. We evaluate LoRSA for four-class breast-density classification using DINOv3-Base, with VinDr-Mammo as the source domain and MammosighTR and RSNA as unseen external domains. LoRSA remains competitive on the internal validation set and achieves the best external macro-F1 on both target datasets, improving upon the strongest competing method by 2.15 percentage points on MammosighTR and 3.09 percentage points on RSNA. Weight-matrix analysis further shows that approximately $92\%$ of the energy of each adaptation component lies outside the bilateral singular subspace of the other, indicating that the two components learn largely complementary update directions. These results suggest that organizing adaptation capacity into distinct global and residual paths can improve the external-domain generalization of parameter-efficiently adapted biomedical vision models.

参数高效微调生物医学图像跨域泛化低秩更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。