提出可验证的知识门控任务构造方法,区分公开指令与私有知识依赖。
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

- 将任务指令与包含私有规范的紧凑知识包分离,显式控制知识依赖
- 无知识包时前沿模型通过率从68.0%降至0%,证明知识必要性
- 支持可执行验证和规则校验,适合评估大模型真实知识掌握能力
专业代理任务常依赖公共语料库中缺失的隐含惯例,但现有基准极少控制代理是否可访问这些惯例。本文提出一种知识门控任务构造协议,将任务指令与包含私有惯例、参考表和实用算子的紧凑实体分离。通过构建期溯源、字节级相同的指令、泄露审计及可执行见证机制,使对知识包的依赖关系显式且可测。在十五个校准任务中,一个前沿模型配置在拥有知识包时通过率为68.0%,无知识包时为0%;在一项任务中,看似合理但错误的知识包同样导致五次试验全失败。确定性求解器与规则语料库为结构化任务提供精确真值,命名准则级别的评分标准则支持无法由单一可执行预言机验证的输出。基于配置相对校准的筛选保留了七项满足五次试验经验知识门控标准的任务。实验验证了该构造协议的有效性;但未证明保留任务能提升训练后性能。部分任务套件及配套工具已开源至https://github.com/DatagridsAI/Knowledge-Gated-Task-Construction。
原文摘要 · Abstract (English)
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators. Construction-time provenance, byte-identical task instructions across the provided- and withheld-artefact conditions, leak audits, and executable witnesses make dependence on the artefact explicit and testable. Across fifteen calibration tasks, one frontier agent configuration achieves a 68.0% pass rate with the artefact and 0% without it; on one task, a plausible but incorrect artefact also yields 0% across five trials. Deterministic solvers and rule corpora provide exact ground truth for structured tasks, while named criterion-level rubrics support outputs that cannot be checked by a single executable oracle. A configuration-relative calibration screen retains seven tasks satisfying our five-trial empirical knowledge-gating screen. These experiments validate the behavior of the construction protocol; they do not establish that the retained tasks improve post-training. We publicly release part of the task suite and supporting tooling at https://github.com/DatagridsAI/Knowledge-Gated-Task-Construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。