arXiv:2607.08961cs.LGcs.AI2026-07

提出NL-PAC框架,量化LLM标注的歧义风险并给出可验证的最小风险下界。

NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision

  • 用阈值解码定义可接受标签集,构建形式化监督歧义分析框架
  • 证明在盲目标监督下,学习者风险至少为歧义直径的一半,且可被数据独立策略达到
  • 可在冻结模型上验证证书,适合评估LLM标注可靠性与提示设计

大型语言模型越来越多地以自然语言形式提供任务标注、评估与反馈。当规范存在多种解读但监督渠道未揭示具体使用哪一种时,增加标签只能减少抽样误差,无法解决识别问题。本文提出自然语言PAC(NL-PAC)框架,利用固定模型的阈值解码规则定义可接受标签与候选目标。多个标签可接受的概率等于点态可接受目标类的直径;在目标不可见监督下,所有学习者在任意样本量下均面临至少该直径一半的最坏情况风险;该类上的随机最小最大风险可由数据无关策略精确实现。有限样本置信区间使这些量可从保留的未标注输入中认证。在冻结的Qwen 2.5–3B模型审计中,一个预设提示产生正的模型相对证书,而改写提示与精确规则对照组则为零。保留桥接审计发现,提供的候选解读条款不满足传递证书所需的可接受性条件。该保证仅针对已审计模型、提示、阈值与输入分布;推广至人类解释需外部验证。

原文摘要 · Abstract (English)

Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification admits multiple readings but the supervision channel does not reveal which is operative, additional labels reduce sampling error without resolving the resulting identification problem. We introduce Natural Language PAC (NL-PAC), a framework that uses a fixed model's thresholded decoding law to define admissible labels and candidate targets. The probability that multiple labels are admissible equals the diameter of the pointwise-admissible target class, and under target-blind supervision every learner incurs worst-case risk of at least half this diameter, at every sample size; the exact randomized minimax risk over this class is attained by a data-independent strategy. Finite-sample confidence bounds make these quantities certifiable from held-out unlabeled inputs. In a frozen Qwen~2.5--3B audit, one prespecified prompt yields a positive model-relative certificate, whereas a paraphrase and exact-rule controls yield zero. A held-out bridge audit finds that supplied candidate reading clauses fail the admissibility condition needed to transfer the certificate to coherent readings. The guarantee is specific to the audited model, prompt, threshold, and input distribution; extending it to human interpretations requires external validation.

LLM标注风险下界可验证性提示设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。