arXiv:2606.04321cs.AI2026-06

让AI逐步获得自主权,靠人类实证批准,确保安全可控。

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

论文配图:The Digital Apprentice: A Framework for Human-Directed Agentic AI Development
图 1 · 摘自论文原文
  • AI通过人类指导逐步获得权限,每步需实证支持
  • 运行中自动纠正偏差,修正数据可转为学习依据
  • 适合需要长期信任的高风险专业场景

当前自主智能体面临根本矛盾:过度人工监管限制规模,完全自治又难追责。本文提出「数字学徒」框架,实现可扩展且安全的智能体发展。该框架下,自主权不是默认赋予,而是通过逐项技能认证逐步获得,仅当实证证据支持时才晋升。系统包含三大组件:(1)方法捕获,将专业人士的隐性经验转化为结构化资产;(2)授权机制,自主权提升需明确人工审批;(3)持续对齐,运行中检测并修正偏差,每次修正转化为专属偏好数据。我们在推理阶段实现该框架,并对公开专业语料进行验证。实验表明,在流量变化导致性能下降时,动态调整策略可恢复质量指标。该框架三支柱协同,为可信赖、可扩展的智能体系统提供新路径。

原文摘要 · Abstract (English)

Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture provides the governance infrastructure required for responsible delegation. We present the Digital Apprentice, a framework for scalable, safe AI agency in which autonomy is earned, not assumed. The Digital Apprentice is a developmental learner that internalizes the tacit methodology of a directing human, graduating through per-skill autonomy tiers only when empirical evidence justifies it. The result is an agent that becomes genuinely useful over time while remaining aligned to a specific human's standards. Three architectural components make this possible. (1) Methodology capture, distilling a directing professional's tacit approach into structured assets. (2) Authorization, with autonomy escalation gated by explicit human approval. (3) Continuous alignment, correcting drift at runtime and converting each correction into owned preference data. We instantiate this framework as an inference-time control plane. We mathematically model the quality framework and discuss policies and techniques designed to raise quality. We apply the framework to an open professional corpus, and we show how catching data drift and applying a different technique at runtime recovers degraded quality dimensions under traffic shift. The implication extends beyond any single application. We believe these three pillars, stitched together as a system, form a safer and more viable path to agentic systems that can scale without sacrificing trust.

自主智能体人机协作持续对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。