arXiv:2509.03986cs.CVcs.AI2025-09EMNLP被引 20

研究大模型对提示词的敏感度,发现微小改动可导致准确率差15%。

Promptception: How Sensitive Are Large Multimodal Models to Prompts?

  • 构建61种提示类型框架,系统测试大模型对提示的敏感性
  • 发现闭源模型比开源模型更受提示词影响,最高偏差达15%
  • 提出适配不同模型的提示原则,提升评估公平性

尽管近年来大型多模态模型(LMMs)取得显著进展,但在多项选择题问答(MCQA)任务中,提示词设计仍缺乏深入理解。我们发现,提示词表述和结构的细微变化,可能导致某些模型和提示的准确率波动高达15%。这种不稳定性给模型评估的透明性和公平性带来挑战,因模型常使用精心挑选的提示展示最佳性能。为此,我们提出Promptception,一个系统性评估提示敏感性的框架。该框架包含61种提示类型,涵盖15个类别与6个超类别,针对提示构造的不同方面。我们利用该框架在3个MCQA基准(MMStar、MMU-Pro、MVBench)上评估了10个LMMs,包括轻量级开源模型至GPT-4o和Gemini 1.5 Pro。结果表明,闭源模型对提示语义更敏感,体现更强指令对齐;而开源模型表现更稳定,但难以应对复杂或细微的提示。基于分析,我们提出适配闭源与开源模型的提示原则,以实现更鲁棒、公平的模型评估。

原文摘要 · Abstract (English)

Despite the success of Large Multimodal Models (LMMs) in recent years, prompt design for LMMs in Multiple-Choice Question Answering (MCQA) remains poorly understood. We show that even minor variations in prompt phrasing and structure can lead to accuracy deviations of up to 15% for certain prompts and models. This variability poses a challenge for transparent and fair LMM evaluation, as models often report their best-case performance using carefully selected prompts. To address this, we introduce Promptception, a systematic framework for evaluating prompt sensitivity in LMMs. It consists of 61 prompt types, spanning 15 categories and 6 supercategories, each targeting specific aspects of prompt formulation, and is used to evaluate 10 LMMs ranging from lightweight open-source models to GPT-4o and Gemini 1.5 Pro, across 3 MCQA benchmarks: MMStar, MMMU-Pro, MVBench. Our findings reveal that proprietary models exhibit greater sensitivity to prompt phrasing, reflecting tighter alignment with instruction semantics, while open-source models are steadier but struggle with nuanced and complex phrasing. Based on this analysis, we propose Prompting Principles tailored to proprietary and open-source LMMs, enabling more robust and fair model evaluation.

多模态模型提示工程模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。