新协议让辩论双方在复杂问题上公平对决,防止作弊方用难题困住对手。
Avoiding Obfuscation with Prover-Estimator Debate
- 引入探针-估计器辩论机制,通过递归拆解问题提升判断力
- 在特定稳定假设下,诚实方可用与对手相当的算力获胜
- 解决旧协议中恶意方制造计算陷阱的漏洞,适合高阶AI对齐研究
训练强大人工智能系统以展现期望行为,依赖于在日益复杂的任务上提供准确的人类监督。一种有前景的方法是利用两个竞争性AI在关于给定问题正确解的辩论中放大人类判断力。先前的理论工作为AI辩论提供了复杂性理论形式化,并提出了设计辩论协议的问题,以确保人类判断在尽可能复杂的任务类别上保持正确。递归辩论(其中辩论者将复杂问题分解为更简单的子问题)有望扩大可被准确判断的问题范围。然而,现有的递归辩论协议面临混淆论证问题:不诚实的辩论者可采用计算高效策略,迫使诚实对手解决计算上不可行的问题才能取胜。本文通过一种新递归辩论协议缓解此问题,在某些稳定性假设下,确保诚实辩论者可采用与其对手计算效率相当的策略获胜。
原文摘要 · Abstract (English)
Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given problem. Prior theoretical work has provided a complexity-theoretic formalization of AI debate, and posed the problem of designing protocols for AI debate that guarantee the correctness of human judgements for as complex a class of problems as possible. Recursive debates, in which debaters decompose a complex problem into simpler subproblems, hold promise for growing the class of problems that can be accurately judged in a debate. However, existing protocols for recursive debate run into the obfuscated arguments problem: a dishonest debater can use a computationally efficient strategy that forces an honest opponent to solve a computationally intractable problem to win. We mitigate this problem with a new recursive debate protocol that, under certain stability assumptions, ensures that an honest debater can win with a strategy requiring computational efficiency comparable to their opponent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。