arXiv:2608.08282cs.LG2026-08

提出精确重用无效证证书的智能体推理方法,提升约束条件下的生成准确性。

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

  • 通过冻结状态-延续模式库,实现跨历史路径的无效性证证书复用。
  • 在极低有效概率下仍保持条件分布精度达10^-16,远超传统方法误差97%。
  • 适合对严格约束条件有高精度需求的复杂决策类大模型应用。

使用工具的语言模型智能体面临随观测和先前动作而变化的约束。本文研究在硬状态验证器条件下对模型分布进行精确采样,并跨历史重用无效性证书。Stateful CARS 在每次尝试中冻结一组无瑕疵的状态-延续模式库,并移除所有包含在匹配抽象状态处被认证为无效的延续轨迹。通过精确的残差Doob变换从剩余提案中采样。我们给出可检验的未来有效性双向模拟条件,证明了模式的正确性、自适应精确性、独立同分布输出、几乎必然终止、单调接受率及压缩不变性,并将计算量定义为可达全历史乘积状态数。对于依赖历史的语言模型,该数量可能呈指数增长;因此所提方法不保证通用有限前缀树可扩展性。在可枚举工作流上,其分析定律在有效概率为6×10^-8时与真实条件分布偏差仅10^-16,而状态感知局部解码误差高达0.97。与基准对比显示:观察键式官方CARS在采样步数上更优(根/状态比0.942 [0.934,0.951]),Qwen比较结果无显著差异(0.99 [0.90,1.08])。跨历史迁移仅在内部匹配键消融中带来1.27倍收益。证据支持基于模式的精确条件化,而非系统层面优于CARS。

原文摘要 · Abstract (English)

Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model distribution conditioned on a hard stateful validator while reusing invalidity certificates across histories. Stateful CARS freezes a bank of sound state--continuation schemas within each attempt and removes every trajectory containing a certified continuation at a matching abstract state. An exact residual Doob transform samples from the resulting proposal. We give a checkable future-validity bisimulation condition, prove schema soundness, adaptive exactness, i.i.d.\ outputs, almost-sure termination, monotone acceptance, and compression invariance, and characterize computation by the number of reachable full-history product states. This number can be exponential for a history-dependent language model; the evaluated method therefore makes no generic finite-trie scalability claim. On enumerable workflows, its analytic law matches the valid conditional to $10^{-16}$ at validity probability $6\times10^{-8}$, whereas state-aware local decoding can be $0.97$ away. A matched comparison is negative: observation-keyed official CARS is cheaper in sampler steps (root/Stateful ratio $0.942$ $[0.934,0.951]$), and the Qwen comparison is null ($0.99$ $[0.90,1.08]$). Cross-history transfer helps only in an internal matched-key ablation ($1.27\times$). Thus the evidence supports exact schema-induced conditioning, not a systems advantage over CARS.

大模型推理约束生成精确采样智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。