arXiv:2503.04199cs.CVcs.AI2025-03被引 1

用大模型融合可见光与热成像,让文本提示直接参与分割。

MASTER: Multimodal Segmentation with Text Prompts

  • 双路径提取多模态特征,大模型生成可学习的融合代码本。
  • 在多个自动驾驶场景中表现优异,支持复杂文本查询。
  • 结构简单适应性强,适合需要文本引导的视觉任务。

RGB-热成像融合是应对复杂场景下恶劣天气和光照条件的潜在解决方案。然而,现有研究多聚焦于设计复杂的多模态融合模块。随着大语言模型(LLMs)的广泛应用,自然语言中的有价值信息得以更高效提取。为此,我们旨在利用大语言模型的优势,设计一种结构简洁、高度可适应的多模态融合模型架构。提出多模态分割文本提示框架(MASTER),将大语言模型融入RGB-热成像数据的融合过程,支持复杂文本查询参与融合。模型采用双路径结构分别提取图像不同模态的信息,并以大语言模型为核心,从RGB图像、热成像和文本信息中生成可学习的代码本令牌。通过轻量级图像解码器获得语义分割结果。MASTER在多个自动驾驶基准测试中表现卓越,展现出良好性能。

原文摘要 · Abstract (English)

RGB-Thermal fusion is a potential solution for various weather and light conditions in challenging scenarios. However, plenty of studies focus on designing complex modules to fuse different modalities. With the widespread application of large language models (LLMs), valuable information can be more effectively extracted from natural language. Therefore, we aim to leverage the advantages of large language models to design a structurally simple and highly adaptable multimodal fusion model architecture. We proposed MultimodAl Segmentation with TExt PRompts (MASTER) architecture, which integrates LLM into the fusion of RGB-Thermal multimodal data and allows complex query text to participate in the fusion process. Our model utilizes a dual-path structure to extract information from different modalities of images. Additionally, we employ LLM as the core module for multimodal fusion, enabling the model to generate learnable codebook tokens from RGB, thermal images, and textual information. A lightweight image decoder is used to obtain semantic segmentation results. The proposed MASTER performs exceptionally well in benchmark tests across various automated driving scenarios, yielding promising results.

多模态分割大模型文本提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。