arXiv:2603.05293cs.LGcs.CL2026-03被引 2

通过知识差异几何分析,揭示辩论在什么情况下比单模型反馈更有效。

Knowledge Divergence and the Value of Debate for Scalable Oversight

  • 用表示空间主角衡量模型间知识差异,推导出辩论优势的精确公式。
  • 当模型知识互补时,辩论优势从二次增长转为线性,成为必要手段。
  • 适用于研究安全对齐、对抗性监督及多模型知识融合的学者。

通过参数化辩论的价值,基于模型间知识差异的几何特性进行分析。利用模型表示子空间间的主角,我们证明了辩论优势存在精确闭式解。当模型共享相同训练语料时,辩论退化为类似RLAIF的单智能体方法,可达到相同最优解。当模型知识存在差异时,辩论优势随相位转变从二次阶段(收益微弱)跃升至线性阶段(辩论至关重要)。我们划分了三类知识差异模式(共享、单向、组合),并给出存在性结果:辩论可实现任一模型单独无法达成的结果;同时发现,在组合模式下,过强的对抗激励会导致协作失败,且存在一个清晰阈值区分有效与无效辩论。本工作首次建立辩论与RLAIF的正式联系,为对抗性监督的合理性提供几何基础,并关联到跨模型隐含知识的挖掘问题。

原文摘要 · Abstract (English)

AI safety via debate and reinforcement learning from AI feedback (RLAIF) are both proposed methods for scalable oversight of advanced AI systems, yet no formal framework relates them or characterizes when debate offers an advantage. We analyze this by parameterizing debate's value through the geometry of knowledge divergence between debating models. Using principal angles between models' representation subspaces, we prove that the debate advantage admits an exact closed form. When models share identical training corpora, debate reduces to RLAIF-like where a single-agent method recovers the same optimum. When models possess divergent knowledge, debate advantage scales with a phase transition from quadratic regime (debate offers negligible benefit) to linear regime (debate is essential). We classify three regimes of knowledge divergence (shared, one-sided, and compositional) and provide existence results showing that debate can achieve outcomes inaccessible to either model alone, alongside a negative result showing that sufficiently strong adversarial incentives cause coordination failure in the compositional regime, with a sharp threshold separating effective from ineffective debate. We offer the first formal connection between debate and RLAIF, a geometric foundation for understanding when adversarial oversight protocols are justified, and connection to the problem of eliciting latent knowledge across models with complementary information.

AI安全辩论机制知识融合对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。