提出可审计的测试协议,精准识别大模型推理错误
Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol
- 用有限组件语法生成确定性提示变体
- 经语义与提取审计后才计为模型错误
- 强调应以审计通过率评估提示策略优劣
固定推理基准评估标准提示,但语义有效的呈现变化仍可能改变模型行为。现有提示变异研究缺乏审计,易混入格式错误、提取伪影和不匹配搜索过程。本文提出审计约束的靶向推理评估协议:提示变体由有限组件语法生成,确定性渲染,在固定查询预算下评估,仅在通过语义与提取审计后才计为模型错误。在此协议下,我们实现基于得分的组件自适应采样(CAPS),并与等预算均匀采样对比,使用相同任务库、渲染器、模型接口、解码设置和审计流程。在三个经审计的片段中,协议识别出确认的模型错误提示键,排除了格式与提取伪影;但匹配对比未显示CAPS在审计通过率或唯一提示键发现上优于均匀采样。贡献在于方法论:靶向提示变异可在可复现、可审查、预算匹配的协议下研究,代理引导策略应以审计通过率而非原始不匹配数或选取样本单独判断。
原文摘要 · Abstract (English)
Fixed reasoning benchmarks evaluate canonical prompts, but semantically valid changes in presentation can still change model behavior. Studies of prompt variation can reveal such failures, but without audit they can mix genuine model errors with invalid perturbations, extraction artifacts, and unmatched search procedures. We propose an audit-constrained protocol for targeted reasoning evaluation. Prompt variants are generated from a finite component grammar, rendered deterministically, evaluated under a fixed query budget, and counted as model errors only after semantic and extraction audit. Within this protocol we instantiate Component-Adaptive Prompt Sampling (CAPS), a score-based sampler over prompt components, and compare it with equal-budget uniform component sampling under the same task bank, renderer, model interface, decoding settings, and audit procedure. Across three audited slices, the protocol identifies confirmed model-error prompt keys while excluding formatting and extraction artifacts, but matched comparisons do not show that CAPS improves audited yield or unique prompt-key discovery over uniform sampling. The contribution is methodological: targeted prompt variation can be studied under a reconstructable, reviewable, budget-matched protocol, and proxy-guided policies should be judged by audited yield rather than raw mismatch counts or selected examples alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。