优化提示词可小幅提升大模型对球迷体验评分的预测准确率。
The signal is the ceiling: Measurement limits of LLM-predicted experience ratings from open-ended survey text
- 通过调整提示词,将预测准确率从67%提升至69%
- 模型选择影响有限,5.2版退回到原始水平,4.1-mini下降6个百分点
- 文本语言特征对结果影响远超提示或模型选择,是主要瓶颈
此前研究(Hong, Potteiger, and Zapata 2026)表明,未经优化的 GPT 4.1 提示词可在约 10,000 份来自五支 MLB 球队的赛后调查文本中,以 67% 的概率预测出球迷报告的体验评分,误差在 ±1 范围内。本文测试了提示设计与模型选择对性能的影响。对比四种配置:原始基线提示与适度定制提示,分别搭配 GPT 4.1、4.1-mini 与 5.2 模型。提示定制在 GPT 4.1 上使 ±1 准确率提升约两个百分点(67% → 69%)。但模型更换均导致性能下降:GPT 5.2 回落至基线,GPT 4.1-mini 下降 6 个百分点。两者结合也难以突破输入文本本身的限制——在不同能力配置下,文本语言特征带来的准确率差异超过一个数量级,远超提示或模型选择的影响。天花板由两部分构成:一是模型解读文本的偏差,可通过提示纠正;二是球迷所写内容与真实评分之间的信息差,无法通过工程手段弥补。提示优化仅能触及第一部分,且效果有限且可预测。
原文摘要 · Abstract (English)
An earlier paper (Hong, Potteiger, and Zapata 2026) established that an unoptimized GPT 4.1 prompt predicts fan-reported experience ratings within one point 67% of the time from open-ended survey text. This paper tests the relative impact of prompt design and model selection on that performance. We compared four configurations on approximately 10,000 post-game surveys from five MLB teams: the original baseline prompt and a moderately customized version, crossed with three GPT models (4.1, 4.1-mini, 5.2). Prompt customization added roughly two percentage points of within +/-1 agreement on GPT 4.1 (from 67% to 69%). Both model swaps from that best configuration degraded performance: GPT 5.2 returned to the baseline, and GPT 4.1-mini fell six percentage points below it. Both levers combined were dwarfed by the input itself: across capable configurations, accuracy varied more than an order of magnitude more by the linguistic character of the text than by the choice of prompt or model. The ceiling has two parts. One is a bias in how the model reads text, which prompt design can correct. The other is a difference between what fans write about and what they actually decide, which no engineering can close because the missing information is not in the text. Prompt customization moved the first part; model selection moved neither reliably. The result is not that "prompt engineering helps a little" but that prompt engineering helps in a specific and predictable way, on the part of the ceiling it can reach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。