通过标准化代码降低编程熵,实现百倍成本缩减。
No Accidental Software Agent First Canonical Code for Human Code Entropy Reduction and 30 to 500 times Lower Frontier Model Requirements
- 提出代理优先的规范代码框架,用行为等价性压缩冗余编码。
- 初步实验显示可学习超6万条规范路径,抑制错误语言标记。
- 适合关注代码可验证性与成本优化的研发团队。
前沿编程模型需耗费大量算力学习人类代码库中的偶然熵,这些库虽包含测试、故障、迁移等宝贵信号,但被框架迭代、命名漂移、生成歧义等因素干扰。本文提出代理优先的规范代码体系,将常规软件重写为行为特征图谱、类型化变更代数、证明路径、约束编辑语法、语义补丁单元、运行时负向记忆和带证明的变更对象。核心假设是:通过声明式断言对软件进行行为等价商化,可将等价编码压缩为受控代表,附带显式证据与证明责任。目标是实现每次经验证的正确变更的摊销成本,涵盖源码、上下文、推理、工具、验证、安全、溯源、评审、失败循环、缺陷及制造成本。报告的降幅为假设,非实测前沿结果。所提极限为‘无意外边界’——偶发性可消除至残余创新、证据、治理、风险与未来选项主导。对支持的常规产品分布,可达成近100倍全成本削减,非普适保证。对Qwen2.5-Coder-14B的预实验显示,64,088条规范轨迹可学习,且抑制测试中禁用语言标记,但未验证行为保持、扩展经济性或验证变更成本。贡献在于一个可证伪的程序,聚焦最小功能描述长度与已验证变更成本。
原文摘要 · Abstract (English)
Frontier coding models may spend substantial capacity learning not only program behavior, but also accidental entropy in human repositories. Such repositories contain valuable signals: tests, incidents, migrations, edge cases, product judgment, and operational history. These signals are entangled with framework churn, naming drift, generated-source ambiguity, dependency rituals, CI dialects, weak proof routes, and human-oriented review customs. We propose agent-first canonical code, a proof-carrying substrate that rewrites routine product software into canonical behavior profiles, typed change algebra, proof lanes, constrained edit grammars, semantic patch cells, runtime negative memory, and proof-carrying change objects. The core hypothesis is that quotienting software by behavior equivalence under a declared oracle can collapse equivalent encodings into governed representatives with explicit evidence and proof obligations. The endpoint is amortized cost per verified correct change, including source, context, reasoning, tools, verification, security, provenance, review, failed loops, defects, and foundry cost under a common oracle. Reported reduction bands are hypotheses, not measured frontier results. The proposed limit is a No-Accident Horizon: removable accident decreases until residual novelty, evidence, governance, risk, and future optionality dominate. For supported routine-product distributions, this gives a defensible planning target near 100-fold all-in cost reduction, not a guarantee for all software. Preliminary QLoRA experiments on Qwen2.5-Coder-14B show that 64,088 canonical trajectories are learnable and suppress tested forbidden-language markers, but do not establish behavior preservation, scaling economics, or verified-change cost. The contribution is a falsifiable program centered on minimum functional description length and verified-change cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。