构建可持续演进的AI原生软件生产系统,通过契约驱动与对抗验证保障可靠性。
Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report
- 将需求转化为显式契约,由专业化AI代理分步执行任务。
- 在17个功能、数周部署中发现契约不全与验证边界问题。
- 适合需要长期运维的AI服务型软件场景,如小型企业CTO服务。
AI原生软件开发常局限于单个模型、提示词或生成产物的评估,但在生产环境中,软件需在多种运行场景和长期周期内持续交付、验证、部署、维护与迭代。本文提出一种元工程框架:将产品与运营需求转化为显式契约,经角色专精的AI代理处理,通过独立且对抗性验证,并基于结构化失败分类与外层校准实现自我优化。该框架适用于软件交付非一次性项目而是持续运营职能的场景。在面向小服务企业的CTO即服务应用中,系统持续管理网站、预订流程、支付系统、后台自动化及AI代理接口等技术基础设施。架构包含两阶段契约编译、带专业记录的持久化Markdown记忆、基于注意力与独立性的双重验证、四向失败仲裁器及外层校准机制。早期生产部署覆盖17项功能,数周时间,其中一次内置支付案例揭示了契约不完整与验证边界缺陷,直接推动框架改进。贡献在于实现了可度量、可扩展的可信验证架构,使AI原生软件服务具备可审计、可进化的能力。
原文摘要 · Abstract (English)
AI-native software development is often evaluated at the level of individual models, prompts, or generated artifacts. This framing is insufficient for production environments where software must be continuously produced, verified, deployed, maintained, and adapted across many operational contexts and long time horizons. We present a meta-engineering harness: a software-production architecture that transforms operational and product feature requirements into explicit contracts, routes work through role-specialized AI agents, performs independent and adversarial verification, and continuously improves itself through structured failure classification and outer-loop calibration. The harness is designed for settings in which software delivery is not a one-time project but an ongoing operating function. In our motivating application, CTO-as-a-service for small service firms, the system manages websites, booking flows, payment systems, backoffice workflow automations, and AI-agent interfaces as continuously evolving technical infrastructure rather than one-off deliverables. We describe the layered architecture, including two-pass contract compilation, persistent markdown memory with specialization records, attention-based and independence-based verifications, a four-way failure arbiter, and outer-loop calibration. We report results from an early production deployment spanning 17 features over several weeks, including a detailed in-app payments case study that revealed contract incompleteness and verification-boundary issues. These observations directly drove targeted improvements to the harness. The contribution is an implemented, measurable, and extensible verification architecture for making AI-native service-as-a-software production reliable, auditable, and improvable over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。