arXiv:2507.14686cs.CV2025-07被引 2

用大模型知识蒸馏,让小模型更好识别没见过的场景。

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

  • 通过多模态提示蒸馏,从大模型迁移语义与视觉感知知识。
  • 在已知、罕见和未知场景上均提升识别准确率,罕见场景误差降低23%。
  • 适合需要低资源部署且泛化能力强的开放词汇场景识别任务。

近期多模态大语言模型(MLLM)具备强零样本能力,但在复杂场景识别任务中表现不佳,且难以在边缘设备部署。传统场景识别模型泛化能力弱,难以识别未见和稀有情况。本文提出将教师级MLLM的知识迁移到小型场景识别模型中,实现开放词汇场景识别(Ov-GSR)。为此,我们设计了多模态交互提示蒸馏(MIPD)框架:首先利用基于LLM的判断推理生成器(JRG)构建富含上下文语义的正负视觉片段与凝视推理;再引入场景感知与实例感知提示,通过负向引导多模态提示对齐(NMPA)模块,将推理与视觉信息有效对齐,捕获整体与感知层面的多模态知识;最后将对齐后的知识蒸馏至学生模型,显著增强其泛化能力,提升对未见场景的理解,并缓解稀有情况下的预测偏差。我们在优化后的Ov-SWiG数据集上验证,对学生模型在已知、罕见及未见场景上的性能均有显著提升,并在HICO-DET数据集上进一步证明其对未见目标检测的优越性。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edge device deployment. Meanwhile, conventional GSR models often lack generalization ability, falling short in recognizing unseen and rare situations. In this paper, we exploit transferring knowledge from a teacher MLLM to a small GSR model to enhance its generalization and zero-shot abilities, thereby introducing the task of Open-vocabulary Grounded Situation Recognition (Ov-GSR). To achieve this, we propose Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations. Specifically, the MIPD framework first leverages the LLM-based Judgmental Rationales Generator (JRG) to construct positive and negative glimpse and gaze rationales enriched with contextual semantic information. The proposed scene-aware and instance-perception prompts are then introduced to align rationales with visual information from the MLLM teacher via the Negative-Guided Multimodal Prompting Alignment (NMPA) module, effectively capturing holistic and perceptual multimodal knowledge. Finally, the aligned multimodal knowledge is distilled into the student Ov-GSR model, providing a stronger foundation for generalization that enhances situation understanding, bridges the gap between seen and unseen scenarios, and mitigates prediction bias in rare cases. We evaluate MIPD on the refined Ov-SWiG dataset, achieving superior performance on seen, rare, and unseen situations, and further demonstrate improved unseen detection on the HICO-DET dataset.

场景识别知识蒸馏开放词汇多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。