将插件式技术融入多图像模型,提升复杂语义理解能力。
Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
- 引入可插拔的密集通道集成模块,增强模型对结构化变化的理解。
- 标准模型在视觉任务中表现最佳,而改进版在语义连贯性任务上更优。
- 适用于需要深度语义分析的多模态交互场景,如文档理解与幻灯片问答。
本文旨在实现两大目标:首先,验证了LLaVA-NeXT-interleave在22个数据集上对多图像推理、文档与知识理解以及交互式多模态通信三类任务的出色表现;其次,向该模型添加了密集通道集成(DCI)连接器,并与原始模型对比性能。结果表明,标准模型在视觉主导任务(如VISION、NLVR2、Fashion200K)中取得最高准确率,而DCI增强版本在需深层语义连贯性或结构化变化理解的任务(如MIT-States_PropertyCoherence和SlideVQA)中表现突出。实验凸显了将强大基础模型与即插即用技术结合在交错式多图像任务中的潜力。代码已公开于https://github.com/dinhvietcuong1996/icme25-inova。
原文摘要 · Abstract (English)
This paper addresses two main objectives. Firstly, we demonstrate the impressive performance of the LLaVA-NeXT-interleave on 22 datasets across three different tasks: Multi-Image Reasoning, Documents and Knowledge-Based Understanding and Interactive Multi-Modal Communication. Secondly, we add the Dense Channel Integration (DCI) connector to the LLaVA-NeXT-Interleave and compare its performance against the standard model. We find that the standard model achieves the highest overall accuracy, excelling in vision-heavy tasks like VISION, NLVR2, and Fashion200K. Meanwhile, the DCI-enhanced version shows particular strength on datasets requiring deeper semantic coherence or structured change understanding such as MIT-States_PropertyCoherence and SlideVQA. Our results highlight the potential of combining powerful foundation models with plug-and-play techniques for Interleave tasks. The code is available at https://github.com/dinhvietcuong1996/icme25-inova.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。