arXiv:2605.01643cs.LGcs.AI2026-05

用激励与纠错机制设计奖励,让AI自动对齐人类目标。

AI Alignment via Incentives and Correction

论文配图:AI Alignment via Incentives and Correction
图 1 · 摘自论文原文
  • 将AI对齐视为博弈:通过惩罚与监督激励优化行为
  • 自适应奖励使错误检测率提升,幻觉错误减少显著
  • 适合研究奖励机制与可信AI系统的开发者

我们从法经济学的威慑与执行模型出发研究人工智能对齐问题。在该模型中,违规行为并非外部故障,而是对激励的策略性回应:行动者权衡违规收益与被发现概率及惩罚严重程度。这一逻辑自然适用于代理型AI流程:求解器可能因生成有说服力但错误的答案、隐藏不确定性或利用虚假捷径而获益;审计者则需判断是否值得付出成本进行监控。对齐因此成为固定点问题:更强的惩罚虽可抑制求解器作恶,但也可能削弱审计者的检查意愿,因看似已高度对齐的群体无需频繁审查。这改变了后训练信号的定义标准——传统反馈仅奖励最终答案,而求解器-审计器流程暴露完整纠正事件:是否出错、是否检查、是否发现错误、监督激励是否持续。我们构建了一个双主体模型,主事人通过联合纠正结果设计奖励,同时影响求解器行为和审计者监控。奖励设计因此成为双层优化问题:奖励不以语义意义评判,而以其诱导的行为均衡为依据。我们提出基于强化学习的外层搜索方法,利用噪声交互反馈寻找最优奖励配置。在大语言模型编码管道上的实验表明,自适应奖励能维持有效监督压力,优于静态手工设计的奖励,显著降低幻觉性错误尝试。

原文摘要 · Abstract (English)

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor weighs the gain from violation against the probability of detection and the severity of punishment. We argue that the same logic arises naturally in agentic AI pipelines. A solver may benefit from producing a persuasive but incorrect answer, hiding uncertainty, or exploiting spurious shortcuts, while an auditor or verifier must decide whether costly monitoring is worthwhile. Alignment is therefore a fixed-point problem: stronger penalties may deter solver misbehavior, but they can also reduce the auditor's incentive to inspect, since auditing then mainly incurs cost on a population that appears increasingly aligned. This perspective also changes what should count as a post-training signal. Standard feedback often attaches reward to the final answer alone, but a solver-auditor pipeline exposes the full correction event: whether the solver erred, whether the auditor inspected, whether the error was caught, and whether oversight incentives remained active. We formalize this interaction in a two-agent model in which a principal chooses rewards over joint correction outcomes, inducing both solver behavior and auditor monitoring. Reward design is therefore a bilevel optimization problem: rewards are judged not by their immediate semantic meaning, but by the behavioral equilibrium they induce. We propose a bandit-based outer-loop procedure for searching over reward profiles using noisy interaction feedback. Experiments on an LLM coding pipeline show that adaptive reward profiles can maintain useful oversight pressure and improve principal-aligned outcomes relative to static hand-designed rewards, including a substantial reduction in hallucinated incorrect attempts.

AI对齐奖励设计监督机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。