arXiv:2607.09665cs.AIcs.CL2026-07

研究提示词格式对大模型评测结果的影响,发现格式差异可导致排名颠倒。

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

论文配图:Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking
图 1 · 摘自论文原文
  • 设计受控令牌的评测协议,量化不同提示格式对模型表现的影响。
  • 平均敏感度指数差超30倍,主要由模型合规性失败导致。
  • 强调评测需报告格式敏感性和解析能力,避免结论失真。

提示词包装器仅在格式上不同,却可能使模型得分发生足够变化,从而颠倒排行榜结论。我们基于受控令牌协议研究这一变异,提出两个互补指标:格式敏感度指数(FSI),即包装器选择带来的准确率波动范围;解析敏感度指数(PSI),即答案解析能力的对应波动范围。在覆盖7个问答任务、5种包装器族、4个指令模型(参数量从7B到72B)的14万次OpenRouter生成中,发现平均FSI在不同模型间相差超过30倍,且主要由合规性失败解释。固定效应回归显示,即使控制任务、模型和包装器后,解析能力仍是准确率的强预测因子。我们认为,不报告包装器方差与合规性的准确率统计上是脆弱的,并为评测与结构化输出部署提供实用建议。

原文摘要 · Abstract (English)

Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we give practical recommendations for both benchmarking and structured-output deployments.

大模型评测提示工程格式敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。