arXiv:2608.22229cs.CL2026-08

让政策规则可执行:用搜索方法生成有实证依据的规范条文。

Grounded Normative Rule Generation with Structured Search

论文配图:Grounded Normative Rule Generation with Structured Search
图 1 · 摘自论文原文
  • 用MCMC搜索五槽与或图,分离逻辑结构与语言表达。
  • 在116个场景上规则质量提升至81.0%,超越传统模型。
  • 适合需要可验证政策的合规系统、自动化代理部署。

规范性规则如机构章程和工作政策需兼具人类可读性和对实际环境记录的操作可验证性。然而,当前的语言生成与结构化输出基准主要奖励表面流畅性或模式符合性,导致操作性根基薄弱。这造成严重漏洞:主流语言模型生成看似合理但无法执行的政策,因依赖不可用的数据日志或范围错位。为此,我们提出地面化规范规则合成(GNRS)问题,并引入GNRS-Search框架,利用马尔可夫链蒙特卡洛(MCMC)采样优化一个离散的五槽与或图(AOG)。通过显式解耦中间操作结构与最终文本生成,该方法将可执行可行性与写作风格分离,使规则失效可在表面实现前被定位。我们在覆盖8类场景共116个受控目标的GNRS-Bench,以及包含53个真实衍生政策任务、隐藏来源条款的RealCharter-Bench上评估该方法。GNRS-Search将平均评分质量从68.8%提升至81.0%,在公开的可执行综合指标中排名第一;系统性槽位干预证实性能提升源于稳健的操作逻辑,而非修辞调优。最终,通过将自动规则起草转化为可检查的搜索问题,本工作为在受监管环境中部署可验证、合规就绪的个人代理提供了基础范式。

原文摘要 · Abstract (English)

Normative rules like institutional charters and workplace policies must be both human-readable and operationally verifiable against actual environment records. However, current language generation and structured-output benchmarks primarily reward surface fluency or schema compliance, leaving operational grounding weakly tested. This creates a critical vulnerability where standard language models generate plausible-sounding policies that fail during enforcement because they rely on unavailable data logs or misaligned scopes. To address this challenge, we formalize the problem as Grounded Normative Rule Synthesis (GNRS) and introduce GNRS-Search, a framework that utilizes Markov Chain Monte Carlo (MCMC) sampling to optimize a discrete, five-slot And-Or Graph (AOG). By explicitly decoupling intermediate operational structure from final prose generation, this method isolates executable feasibility from writing style and allows rule failures to be localized prior to surface realization. We evaluate our approach on GNRS-Bench, a benchmark spanning 116 controlled goals across eight scene families, and RealCharter-Bench, which evaluates transfer to 53 real-derived policy tasks with hidden source clauses. GNRS-Search raises average rubric quality from 68.8% to 81.0% and ranks first under a disclosed executable composite metric, while systematic slot interventions confirm that performance gains stem from robust operational logic rather than rhetorical tuning. Ultimately, by transforming automated rule drafting into an inspectable search problem, this work provides a foundational paradigm for deploying verifiable and compliance-ready personal agents within regulated environments.

规则生成可验证性结构搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。