arXiv:2412.20750cs.CV2024-12被引 4

让视觉语言模型读懂非RGB传感器图像,低成本提升跨传感器理解能力。

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

  • 用少量特定数据微调,让模型学会非RGB图像特征。
  • 在多种传感器上表现更优,且无需修改原有模型结构。
  • 适合需要部署于多传感器环境的AI系统开发者。

大规模视觉语言模型(VLMs)在对齐视觉与文本输入方面取得显著进展,但其对非RGB传感器图像独特物理特性的深层理解仍有限。本文重新审视这些局限性,提出一种新型、低成本范式,显著提升传感器图像理解能力,且无需大量训练数据或修改现有VLM架构。具体而言,我们提出传感器感知属性微调(SAFT)与多样负属性优化(DNA),仅利用极少传感器特定数据,即可有效学习非RGB特征,并克服当前VLM固有的RGB中心偏见。此外,我们构建了首个全面公开的基准测试集VS-TDX,用于在多样化真实场景下严格评估VLM的传感器特定理解能力。在多种传感器模态和VLM上的大量实验表明,该方法在资源受限且架构不变条件下,始终表现出更优性能与更强泛化能力。本方法为实现VLM在日益多样化的现实传感器环境中的可扩展部署提供了实用进展。

原文摘要 · Abstract (English)

Large-scale Vision-Language Models (VLMs) have achieved notable progress in aligning visual inputs with text. However, their ability to deeply understand the unique physical properties of non-RGB vision sensor images remains limited. In this paper, we revisit and analyze these limitations and introduce a novel, cost-efficient paradigm that significantly advances sensor image understanding-without requiring extensive training data or any modifications to the existing VLM architectures. Specifically, we propose Sensor-Aware Attributes Fine-Tuning (SAFT) with the Diverse Negative Attributes (DNA) optimization, which leverages minimal sensor-specific data to enable robust learning of non-RGB characteristics and overcome RGB-centric biases inherent in current VLMs. In addition, we present VS-TDX-the first comprehensive, public benchmark designed to rigorously evaluate VLMs' sensor-specific understanding across diverse and realistic scenarios. Through extensive experiments on VLMs and various sensor modalities, we validate that our method consistently delivers superior performance and generalization under resource-constrained and architecture-invariant settings. Our approach provides a practical advance towards scalable deployment of VLMs in increasingly sensor-diverse real-world environments.

视觉语言模型多传感器理解低资源微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。