arXiv:2507.08021cs.CLcs.AI2025-07被引 2

探究图像描述中上下文示例配置对大模型效果的影响

Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis

  • 从示例数量、图像检索和标题分配三方面优化上下文配置
  • 发现示例配置显著影响模型性能,且不同配置导致注意力模式差异
  • 适合研究多模态大模型推理机制与效率优化的研究者

大型模型的演进催生了上下文学习(ICL)能力。在自然语言处理领域,大量研究已证实ICL的有效性。受大语言模型成功的启发,研究人员开发出具备ICL能力的大规模多模态模型(LMMs)。然而,针对多模态ICL的示例配置探索仍处于初级阶段。此外,通过控制上下文示例(ICEs)可高效观察和分析LMM在不同输入下的推理特征。本文对图像描述任务中的多模态上下文学习进行了外部与内部的综合研究。外部方面,从样本数量、图像检索策略和标题分配三个维度系统评估示例配置;采用多种指标进行量化分析并总结关键发现。内部方面,分析典型LMM的注意力特性,提出基于注意力的度量方法以量化模型行为,并开展辅助实验验证注意力驱动的模型加速与压缩可行性。进一步对比相同架构和预训练策略下LMMs的性能差异,从预训练数据特征角度解释其成因。研究揭示了示例配置如何通过外部实验影响模型表现,以及通过内部检测识别出的典型注意力模式,为理解多模态ICL提供了双重视角。本研究提出的内外结合分析方法及新度量工具可推广至更广泛的研究领域。

原文摘要 · Abstract (English)

The evolution of large models has witnessed the emergence of In-Context Learning (ICL) capabilities. In Natural Language Processing (NLP), numerous studies have demonstrated the effectiveness of ICL. Inspired by the success of Large Language Models (LLMs), researchers have developed Large Multimodal Models (LMMs) with ICL capabilities. However, explorations of demonstration configuration for multimodal ICL remain preliminary. Additionally, the controllability of In-Context Examples (ICEs) provides an efficient and cost-effective means to observe and analyze the inference characteristics of LMMs under varying inputs. This paper conducts a comprehensive external and internal investigation of multimodal in-context learning on the image captioning task. Externally, we explore demonstration configuration strategies through three dimensions: shot number, image retrieval, and caption assignment. We employ multiple metrics to systematically and thoroughly evaluate and summarize key findings. Internally, we analyze typical LMM attention characteristics and develop attention-based metrics to quantify model behaviors. We also conduct auxiliary experiments to explore the feasibility of attention-driven model acceleration and compression. We further compare performance variations between LMMs with identical model design and pretraining strategies and explain the differences from the angles of pre-training data features. Our study reveals both how ICEs configuration strategies impact model performance through external experiments and characteristic typical patterns through internal inspection, providing dual perspectives for understanding multimodal ICL in LMMs. Our method of combining external and internal analysis to investigate large models, along with our newly proposed metrics, can be applied to broader research areas.

多模态上下文学习图像描述模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。