让超级智能分解任务,用人类能管的子任务实现动态价值观对齐。
Superalignment with Dynamic Human Values
- 将复杂任务拆解为人类可指导的子任务,实现可扩展对齐。
- 提出部分到整体泛化假说,子任务对齐可推广至完整任务。
- 适合关注超级智能安全与价值对齐的研究者。
对齐的两大核心挑战是可扩展的监督和人类价值观的动态性。尽管递归奖励建模等方法解决了前者,却未同时应对后者。我们提出一种新算法框架的路线图:训练超人类推理模型,将复杂任务分解为仍可接受人类指导的子任务。该方法基于‘部分到整体泛化假说’,即子任务解决方案的对齐可泛化至完整任务的对齐。我们主张需测量这一泛化能力,并提出未来改进方向。
原文摘要 · Abstract (English)
Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。