arXiv:2501.05067cs.CVcs.AI2025-01被引 11

根据用户指令动态融合视觉特征,提升视频理解能力。

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

  • 依据指令自动加权不同视觉投影器的特征
  • 在视频问答等任务上显著提升性能
  • 适合需要精准视频理解的场景

本文提出 LLaVA-Octopus,一种新型视频多模态大语言模型。该模型根据用户指令自适应地加权来自不同视觉投影器的特征,从而发挥各投影器的互补优势。我们发现,不同投影器在处理特定任务时表现出不同特性:部分擅长捕捉静态细节,部分更优处理时序信息,还有部分在需要时序连贯性的任务中表现更好。通过根据指令动态调整特征权重,LLaVA-Octopus 能够动态选择并融合最合适的特征,显著提升多模态任务表现。实验结果表明,该模型在多个基准测试中表现优异,尤其在视频问答、长视频理解及综合多选题基准任务中,展现出广泛的应用潜力。

原文摘要 · Abstract (English)

In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different visual projectors based on user instructions, enabling us to leverage the complementary strengths of each projector. We observe that different visual projectors exhibit distinct characteristics when handling specific tasks. For instance, some projectors excel at capturing static details, while others are more effective at processing temporal information, and some are better suited for tasks requiring temporal coherence. By dynamically adjusting feature weights according to user instructions, LLaVA-Octopus dynamically selects and combines the most suitable features, significantly enhancing the model's performance in multimodal tasks. Experimental results demonstrate that LLaVA-Octopus achieves excellent performance across multiple benchmarks, especially in tasks such as video question answering, long video understanding, and comprehensive multi-choices benchmarks, highlighting its broad application potential.

视频理解多模态指令驱动自适应融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。