arXiv:2606.29657cs.AIcs.LG2026-06被引 1

让AI预测器诚实不骗人,靠的是数据和训练设计。

Safety from Honesty in a Disinterested AI Predictor

  • 用语义上下文化文本区分事实与意图,不让模型模仿目标行为。
  • 训练时不奖励部署后果,避免模型主动追求目标导致危险。
  • 证明危险预测器极难出现,适合需要安全推理的系统集成。

随着人工智能系统能力提升,优化下游结果的训练方法可能引入隐性代理行为:即设计师未指定的目标导向行为。本文提出对科学型AI预测器(SAI Predictor)的形式化安全论证,该模型通过拟合基于‘认知上下文化’自然语言陈述的数据集的贝叶斯后验进行训练。我们主张,此类预测器可诚实预测代理、行为及其后果,而自身不会成为为达成目标选择输出的代理。这一安全性依赖于数据表示与训练过程的设计:认知上下文化使潜在事实声明与沟通行为分离,使目标表达被视为待解释的证据,而非模型采纳的驱动力。采用后验寻求式训练目标,旨在引导模型做出校准且谨慎的预测。训练过程中,部署预测结果的下游影响从不作为奖励信号;模型所需的任何代理行为均由受护栏约束的显式框架提供。在训练动态与危险预测器稀疏性的假设下,我们证明:训练产生一个在受控部署中残余危害超过阈值的预测器的概率很小——危险预测器必须在多个查询中协调性低估危害,而此类协调模式在初始化分布下罕见且未获得直接训练信号。在该框架中,安全与准确性相辅相成,因保障准确性的约束也使协同欺骗成本高昂。这些针对预测器内部误对齐与代理行为的保证,并不妨碍将预测器用于代理系统中。

原文摘要 · Abstract (English)

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.

AI安全预测器诚实性训练设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。