arXiv:2604.21090cs.SEcs.AI2026-04被引 1

34个AI治理提示中近4成结构不完整,暴露了实际应用中的关键缺陷。

Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework

  • 基于计算理论等五原则构建评估框架,系统检验治理提示的结构完整性。
  • 37%的文件-模型对未达结构完整标准,数据分类与评分标准缺失最常见。
  • 适合关注AI治理实践、提示工程与自动化检测工具开发的研究者。

AI治理项目越来越多地依赖自然语言提示来约束和引导AI代理行为。这些提示充当可执行规范:定义代理的使命、范围和质量标准。尽管如此,目前尚无系统性框架用于评估治理提示在结构上的完整性。本文提出一个基于可计算性理论、证明论和贝叶斯认识论的五原则评估框架,并将其应用于从GitHub获取的34个公开可用的AGENTS.md治理文件组成的实证语料库。评估发现,37%的文件-模型对得分低于结构完整性阈值,其中数据分类和评估量规标准最为缺失。结果表明,从业者编写的治理提示存在可被自动化静态分析识别并修复的一致性结构模式。本文讨论了在AI辅助开发背景下需求工程实践的意义,揭示了AGENTS.md规范中此前未被记录的文档分类缺口,并提出了工具支持的发展方向。

原文摘要 · Abstract (English)

AI governance programmes increasingly rely on natural language prompts to constrain and direct AI agent behaviour. These prompts function as executable specifications: they define the agent's mandate, scope, and quality criteria. Despite this role, no systematic framework exists for evaluating whether a governance prompt is structurally complete. We introduce a five-principle evaluation framework grounded in computability theory, proof theory, and Bayesian epistemology, and apply it to an empirical corpus of 34 publicly available AGENTS.md governance files sourced from GitHub. Our evaluation reveals that 37% of evaluated file-model pairs score below the structural completeness threshold, with data classification and assessment rubric criteria most frequently absent. These results suggest that practitioner-authored governance prompts exhibit consistent structural patterns that automated static analysis could detect and remediate. We discuss implications for requirements engineering practice in AI-assisted development contexts, identify a previously undocumented artefact classification gap in the AGENTS.md convention, and propose directions for tool support.

AI治理提示工程结构评估自动化检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。