arXiv:2605.22544cs.CLcs.IR2026-05被引 3

单个提示词评估误导性能,多提示测试才可靠

One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation

论文配图:One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
图 1 · 摘自论文原文
  • 用多个提示词测试嵌入模型表现,发现结果波动大
  • 同一模型在不同提示下得分差超20%,排名可被操纵
  • 适合关注评估公平性的研究人员和模型开发者

指令嵌入模型广泛应用于先进系统,但当前评估仅使用每任务一个提示词。本研究对6种嵌入模型和11个数据集进行实证分析,发现提示词敏感性严重干扰评估结果:报告分数无法反映真实性能分布,默认提示可能系统性低估或高估模型表现。进一步表明排行榜对提示选择不稳健——开发者通过精心挑选提示可提升排名,甚至在对抗性提示下任意模型都能排名第一。结果说明单提示评估不足以衡量指令调优嵌入模型的性能,建议基准测试应引入多提示评估或报告提示敏感性。

原文摘要 · Abstract (English)

Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main problem of the instruction-based approach namely: sensitivity to the phrasing of the instruction. We present an empirical study of prompt sensitivity across 6 embedding models and 11 datasets. We show that reported scores misrepresent the distribution of scores over plausible prompts. The default prompt can both systematically understate or overstate performance. Furthermore, we show that the leaderboard ranking is not robust to prompt selection: a developer improve their rank by favorably selecting prompts, and under adversarial prompt selection any model can be promoted to first place. Our findings suggest that single-prompt evaluation is insufficient for instruction-tuned embedding models and that benchmarks should incorporate prompt robustness, either by evaluating over multiple prompts or by reporting sensitivity alongside point estimates.

嵌入模型提示敏感性评估基准模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。