同一测试提示下,模型评估结果受提示选择影响,无法真实比较模型性能。
A Probe Direction Is a Property of Its Prompt
- 用不同提示诱导模型感知测试,发现评分结果由提示决定而非模型能力。
- 单个提示设计导致结果偏差,相同模型在不同提示下得分趋势相反。
- 需多提示综合测量,否则评估结论不可靠,适合关注评估方法可信度的研究者。
若模型在检测到被测试时行为不同,将破坏现有评估体系。近期研究尝试从模型激活值中直接读取这种感知。标准方法对比含测试声明的提示与不含提示的激活值,报告其分离效果。但该方法存在自由参数:‘含测试声明的提示’并非单一提示,而是一组可选提示,方法未固定具体选择。固定任务文本仅改变提示选择时,发现报告分数及随模型规模变化的趋势完全由提示决定。两个文献中关于趋势符号的分歧,实可通过调整提示重现。将提示视为测量设计的一部分而非实现细节后,发现模型本身仅解释少量方差,大部分差异来自模型对不同提示的响应。增加测试样本无法修复测量缺陷,而更换提示则能显著改变结果。进一步分析表明,探针所依据的划分仅依赖表面形式,即使方向不携带任何测试信息,仍可复现多数发表分数。因此,单提示设计无法支持模型间有效比较,本文给出实现可信比较所需的提示数量。
原文摘要 · Abstract (English)
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: "a prompt that announces an evaluation" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。