arXiv:2602.08211cs.CV2026-02被引 1

不微调模型,用图文链式提示提升视觉理解能力

Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension

  • 用多轮图文上下文构建推理链,无需训练即可增强模型
  • 在多个数据集上提升5%至30%的定位准确率
  • 适合想快速优化大模型视觉能力的研究者

给定文本描述,参照表达理解(REC)任务需定位图像中被提及的对象。多模态大语言模型(MLLMs)通过扩大模型规模和训练数据,在REC基准上已取得高准确率。此外,利用思维链(Chain-of-Thought)与工具调用等技术,为模型提供额外的视觉或文本上下文,可进一步提升性能。本文分析了多种通过工具调用提供额外视觉与文本上下文的技术对MLLM在REC任务上的影响,并提出一种无需训练的改进框架——链式标题(Chain-of-Caption)。我们在RefCOCO/RefCOCOg/RefCOCO+和Ref-L4数据集上进行实验,结果表明,单独使用文本或视觉上下文即可在不微调的情况下提升REC性能。结合多种上下文后,本框架在不同交并比(IoU)阈值下相较基线模型准确率提升5%至30%。

原文摘要 · Abstract (English)

Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (MLLMs) have achieved high accuracy on REC benchmarks through scaling up the model size and training data. Moreover, the performance of MLLMs can be further improved using techniques such as Chain-of-Thought and tool use, which provides additional visual or textual context to the model. In this paper, we analyse the effect of various techniques for providing additional visual and textual context via tool use to the MLLM and its effect on the REC task. Furthermore, we propose a training-free framework named Chain-of-Caption to improve the REC performance of MLLMs. We perform experiments on RefCOCO/RefCOCOg/RefCOCO+ and Ref-L4 datasets and show that individual textual or visual context can improve the REC performance without any fine-tuning. By combining multiple contexts, our training-free framework shows between 5% to 30% performance gain over the baseline model on accuracy at various Intersection over Union (IoU) thresholds.

多模态视觉理解无训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。