让AI理解图像中物体的文本上下文,提升多模态交互精准度。
Region-Level Context-Aware Multimodal Understanding
- 通过引入物体框坐标和文本信息,实现区域级上下文感知的多模态理解。
- 在多个任务上显著超越基线模型,尤其在个性化对话和多模态检索中表现突出。
- 开源数据集、评测基准与模型,适合多模态理解与个性化应用研究者使用。
尽管多模态大模型取得进展,但现有研究主要关注通用视觉理解,忽视了将对象关联文本上下文以实现更精准的区域级上下文感知多模态理解(RCMU)的能力。为此,我们首次提出RCMU任务,要求模型在响应用户指令时融合图像内容与区域/对象的文本信息。为赋予多模态大模型该能力,我们提出区域级上下文感知视觉指令微调(RCVIT),将物体信息注入模型输入,并利用边界框坐标实现视觉与文本信息的有效对齐。针对数据缺失问题,我们构建了大规模的RCMU数据集,涵盖多种RCMU任务。同时提出RC&P-Bench评测基准,可全面评估模型在RCMU及多模态个性化理解任务中的表现。此外,设计无参考评价指标,实现对区域级上下文感知图像描述的细粒度评估。基于RCMU数据集对Qwen2-VL进行RCVIT微调,得到RC-Qwen2-VL模型。实验表明,该模型在多项RCMU任务中表现优异,并成功应用于多模态RAG与个性化对话。相关数据、模型与基准已开源。
原文摘要 · Abstract (English)
Despite significant progress, existing research on Multimodal Large Language Models (MLLMs) mainly focuses on general visual understanding, overlooking the ability to integrate textual context associated with objects for a more context-aware multimodal understanding -- an ability we refer to as Region-level Context-aware Multimodal Understanding (RCMU). To address this limitation, we first formulate the RCMU task, which requires models to respond to user instructions by integrating both image content and textual information of regions or objects. To equip MLLMs with RCMU capabilities, we propose Region-level Context-aware Visual Instruction Tuning (RCVIT), which incorporates object information into the model input and enables the model to utilize bounding box coordinates to effectively associate objects' visual content with their textual information. To address the lack of datasets, we introduce the RCMU dataset, a large-scale visual instruction tuning dataset that covers multiple RCMU tasks. We also propose RC\&P-Bench, a comprehensive benchmark that can evaluate the performance of MLLMs in RCMU and multimodal personalized understanding tasks. Additionally, we propose a reference-free evaluation metric to perform a comprehensive and fine-grained evaluation of the region-level context-aware image descriptions. By performing RCVIT on Qwen2-VL models with the RCMU dataset, we developed RC-Qwen2-VL models. Experimental results indicate that RC-Qwen2-VL models not only achieve outstanding performance on multiple RCMU tasks but also demonstrate successful applications in multimodal RAG and personalized conversation. Our data, model and benchmark are available at https://github.com/hongliang-wei/RC-MLLM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。