arXiv:2409.16084cs.CV2024-09AAAI被引 17

构建首个伪装物体多模态数据集,提升大模型在隐蔽场景的识别能力。

MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object Scenarios

  • 构建包含1.1万+图文对的MM-CamObj数据集,分对齐与指令两部分
  • 提出CamObj-Llava模型,在600张图7项任务上比GPT-4o高25.84%准确率
  • 设计六阶段课程学习策略,助力模型有效掌握伪装物体知识

大型视觉语言模型(LVLMs)在多个应用中取得显著进展,但在包含伪装物体的复杂场景中仍面临挑战,主要因训练数据缺乏相关样本。为此,我们首次构建了MM-CamObj数据集,包含两个子集:用于视觉-语言对齐与知识注入的CamObj-Align(11,363个图像-文本对),以及用于微调模型指令跟随能力的CamObj-Instruct(11,363张图像、68,849条多样化对话)。基于该数据集,我们提出专为伪装场景设计的CamObj-Llava模型,并引入六种模式的课程学习策略,以增强模型对伪装物体和场景的知识获取能力。此外,我们构建了CamObj-Bench评估基准,包含600张图像、7项任务及9,449个问题,用于评测现有LVLM在伪装场景下的理解、识别、定位与计数能力。在该基准上对CamObj-Llava、8个开源及3个闭源模型进行实验,结果表明其在7项任务中的4项优于GPT-4o达25.84%。代码与数据集将公开于https://github.com/JCruan519/MM-CamObj。

原文摘要 · Abstract (English)

Large visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving camouflaged objects. This is primarily due to the lack of samples related to camouflaged scenes in the training dataset. To mitigate this issue, we construct the MM-CamObj dataset for the first time, comprising two subsets: CamObj-Align and CamObj-Instruct. Specifically, CamObj-Align contains 11,363 image-text pairs, and it is designed for VL alignment and injecting rich knowledge of camouflaged scenes into LVLMs. CamObj-Instruct is collected for fine-tuning the LVLMs with improved instruction-following capabilities, and it includes 11,363 images and 68,849 conversations with diverse instructions. Based on the MM-CamObj dataset, we propose the CamObj-Llava, an LVLM specifically designed for addressing tasks in camouflaged scenes. To facilitate our model's effective acquisition of knowledge about camouflaged objects and scenes, we introduce a curriculum learning strategy with six distinct modes. Additionally, we construct the CamObj-Bench to evaluate the existing LVLMs' capabilities of understanding, recognition, localization and count in camouflage scenes. This benchmark includes 600 images and 7 tasks, with a total of 9,449 questions. Extensive experiments are conducted on the CamObj-Bench with CamObj-Llava, 8 existing open-source and 3 closed-source LVLMs. Surprisingly, the results indicate that our model achieves a 25.84% improvement in 4 out of 7 tasks compared to GPT-4o. Code and datasets will be available at https://github.com/JCruan519/MM-CamObj.

多模态伪装检测视觉语言模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。