arXiv:2607.14109cs.CLcs.AI2026-07被引 1

复杂提示技巧未必更好,简单提示反而更稳。

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

论文配图:Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
图 1 · 摘自论文原文
  • 测试8种提示方法在10个数据集上的表现,发现基础提示最可靠。
  • 复杂提示多数表现差,最高落后31个百分点,仅两种微增3个百分点。
  • 适合关注模型真实能力而非提示优化的研究者与实践者。

评估大语言模型(LLM)能力与构建多选题问答(MCQA)稳健方案仍是自然语言理解的核心挑战。随着LLM迅速增多,人们普遍认为更复杂的提示技巧能带来更好性能。尽管多项研究声称复杂提示效果更优,但缺乏全面评估。本文通过在10个MCQA数据集上对8种提示技术进行大规模实证研究,涵盖27种模型配置,共评估约4,300道独特题目超过43万次。结果揭示显著悖论:基础提示在多个基准上始终优于复杂推理技巧。仅极简的专家角色提示(CoT-Expert)和归纳角色提示(CoT-Inductive)带来约3个百分点的统计显著提升,其余所有复杂方法均持平或显著落后,最高差距达31~pp(Self-Analogical)。我们进一步分析三大现象:(1) Qwen3-30B-A3B-Thinking-2507 在埃洛评分中意外领先;(2) 不同思维预算模型间存在性能-效率权衡,最优配置依赖模型特性;(3) 数据集难度差异巨大,60%基准准确率低于70%,最难与最易间差距达47.5~pp,表明仍有巨大提升空间。研究暗示评估社区可能过度复杂化提示工程,真正机会在于模型本身改进而非提示优化。

原文摘要 · Abstract (English)

Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim better performance with more sophisticated prompting techniques, but do not provide a comprehensive evaluation. We address this gap through a comprehensive empirical study of 8 prompting techniques across 10 multiple-choice question answering (MCQA) datasets, encompassing 27 model configurations and roughly 4,300 unique questions evaluated more than 430,000 times. Our findings reveal a striking paradox that baseline prompting consistently outperforms complex reasoning techniques on various benchmarks. Only minimal expert and inductive role framing (CoT-Expert and CoT-Inductive) yields a small but statistically significant $\sim$3 percentage-point (pp) gain over baseline whereas every other elaborate technique we tested matches or under-performs it, often by large margins (up to 31~pp for Self-Analogical). We further investigate three critical phenomena: (1) the unexpected victory of Qwen3-30B-A3B-Thinking-2507 in Elo ratings, (2) the performance-efficiency trade-offs across model variants with different thinking budgets, revealing model-dependent optimal configurations, and (3) the substantial variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 47.5~pp spread from easiest to hardest, indicating considerable room for model improvement. These results suggest that the LLM evaluation community may be overcomplicating prompt engineering and that substantial performance gaps remain across diverse benchmarks, offering opportunities for genuine model improvements rather than prompt optimization.

提示工程大模型评估多选题问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。