arXiv:2601.08882cs.CVcs.AI2026-01

用流形约束优化压缩视觉变压器,实现边缘设备高效地理空间模型

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

  • 基于流形约束优化框架DLRT,对齐下游任务目标进行结构化参数压缩
  • 在多个地理空间基准上实现显著参数减少,精度损失极小
  • 适合需在资源受限设备上部署高性能地理空间模型的场景

在资源受限的边缘设备上部署地理空间基础模型,需要紧凑的架构以保持高下游性能。然而,其庞大的参数量及压缩常导致的精度下降限制了实际应用。本文利用流形约束优化框架DLRT,在迁移学习中压缩基于视觉变换器的地理空间基础模型。通过强制参数化结构与下游目标对齐,该方法在保持任务特定精度的同时实现强压缩。实验表明,该方法优于现有的低秩方法(如LoRA)。在多种地理空间基准上的测试证实,实现了显著的参数减少且精度损失最小,使高性能、可本地部署的地理空间模型成为可能。

原文摘要 · Abstract (English)

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induced by compression limit practical adoption. In this work, we leverage manifold-constrained optimization framework DLRT to compress large vision transformer-based geospatial foundation models during transfer learning. By enforcing structured low-dimensional parameterizations aligned with downstream objectives, this approach achieves strong compression while preserving task-specific accuracy. We show that the method outperforms of-the-shelf low-rank methods as LoRA. Experiments on diverse geospatial benchmarks confirm substantial parameter reduction with minimal accuracy loss, enabling high-performing, on-device geospatial models.

视觉变换器模型压缩地理空间边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。