arXiv:2503.06003cs.CV2025-03

用频域特征+低秩适配,让视觉语言模型更高效鲁棒

Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models

  • 融合傅里叶变换与低秩适配,提升特征表达能力
  • 在含高斯噪声数据上性能接近CLIP和SigLIP
  • 适合无人机等实时弱光场景的视觉理解任务

情境感知应用高度依赖对视觉与文本数据的实时处理以提供可操作洞察。视觉语言模型(VLMs)通过连接视觉输入与自然语言描述,成为解析复杂环境的关键工具。然而,这些模型在真实环境中常面临计算挑战。本研究提出一种新型视觉语言模型框架,结合频域变换与低秩适配(LoRA),增强特征提取、可扩展性与效率。不同于仅依赖空间域表示的传统VLM,该方法采用基于离散傅里叶变换(DFT)的低秩特征,同时保留预训练空间权重,显著提升在噪声或低光照条件下的鲁棒性。我们在含不同高斯噪声水平的基准数据集上评估了模型在图像描述生成与视觉问答(VQA)任务中的表现。定量结果显示,模型性能与当前顶尖VLM如CLIP ViT-L/14和SigLIP相当。定性分析表明,模型对真实世界图像(由安装于无人地面车辆的RealSense相机采集)能生成更详细且语境相关性强的响应。

原文摘要 · Abstract (English)

Situational awareness applications rely heavily on real-time processing of visual and textual data to provide actionable insights. Vision language models (VLMs) have become essential tools for interpreting complex environments by connecting visual inputs with natural language descriptions. However, these models often face computational challenges, especially when required to perform efficiently in real environments. This research presents a novel vision language model (VLM) framework that leverages frequency domain transformations and low-rank adaptation (LoRA) to enhance feature extraction, scalability, and efficiency. Unlike traditional VLMs, which rely solely on spatial-domain representations, our approach incorporates Discrete Fourier Transform (DFT) based low-rank features while retaining pretrained spatial weights, enabling robust performance in noisy or low visibility scenarios. We evaluated the proposed model on caption generation and Visual Question Answering (VQA) tasks using benchmark datasets with varying levels of Gaussian noise. Quantitative results demonstrate that our model achieves evaluation metrics comparable to state-of-the-art VLMs, such as CLIP ViT-L/14 and SigLIP. Qualitative analysis further reveals that our model provides more detailed and contextually relevant responses, particularly for real-world images captured by a RealSense camera mounted on an Unmanned Ground Vehicle (UGV).

视觉语言模型频域特征低秩适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。