现有评估方法可能只是在识别格式而非真正理解上下文。
Is Evaluation Awareness Just Format Sensitivity? Limitations of Probe-Based Evidence under Controlled Prompt Structure
- 用受控数据集测试模型对提示格式的敏感性
- 探针信号主要依赖标准提示结构,无法泛化到自由格式
- 适合关注评估可靠性与模型真实理解力的研究者
先前研究通过在线性探针上使用基准提示来证明大语言模型具备评估意识。由于评估上下文通常与提示格式和体裁混杂,无法确定探针信号是反映真实语境还是表面结构。本文通过受控的2x2数据集和诊断性重写,在部分控制提示格式的情况下进行测试。结果表明,探针主要追踪基准-标准结构,无法在独立于语言风格的自由格式提示中保持有效。因此,现有基于探针的方法难以可靠区分评估上下文与结构伪影,削弱了已有结论的证据力度。
原文摘要 · Abstract (English)
Prior work uses linear probes on benchmark prompts as evidence of evaluation awareness in large language models. Because evaluation context is typically entangled with benchmark format and genre, it is unclear whether probe-based signals reflect context or surface structure. We test whether these signals persist under partial control of prompt format using a controlled 2x2 dataset and diagnostic rewrites. We find that probes primarily track benchmark-canonical structure and fail to generalize to free-form prompts independent of linguistic style. Thus, standard probe-based methodologies do not reliably disentangle evaluation context from structural artifacts, limiting the evidential strength of existing results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。