arXiv:2411.10534cs.HCcs.AI2024-11被引 2

用公众意愿指导大模型对齐,实现安全可控的智能对话。

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment

  • 构建从公众意愿到专家规则的对齐链,分步实现模型行为优化。
  • 在心理健康领域验证,模型评估与人类专家一致性达0.841(皮尔逊相关)。
  • 适合关注AI伦理、公共安全的开发者和政策制定者参考。

我们提出一种衡量语言模型(LM)行为与公众意愿对齐程度的方法,可用于微调、在线监督和发布前安全检测。该方法通过‘对齐链’(CoA)生成基于规则的奖励(RBR),将公众意愿转化为规范性目标,再由专家设计实现目标的模型行为规则。我们在三个与心理健康相关的提示领域验证了该方法:通过集体对话与桥接排序,使至少96%±2%的美国公众支持所生成的规范性目标;专家制定的规则构建的RBR,在评估模型响应时与人类专家评分高度一致(皮尔逊相关r=0.841,AUC=0.964)。该方法为衡量模型行为与公众意愿的一致性提供了近似有效指标。

原文摘要 · Abstract (English)

We introduce a method to measure the alignment between public will and language model (LM) behavior that can be applied to fine-tuning, online oversight, and pre-release safety checks. Our `chain of alignment' (CoA) approach produces a rule based reward (RBR) by creating model behavior $\textit{rules}$ aligned to normative $\textit{objectives}$ aligned to $\textit{public will}$. This factoring enables a nonexpert public to directly specify their will through the normative objectives, while expert intelligence is used to figure out rules entailing model behavior that best achieves those objectives. We validate our approach by applying it across three different domains of LM prompts related to mental health. We demonstrate a public input process built on collective dialogues and bridging-based ranking that reliably produces normative objectives supported by at least $96\% \pm 2\%$ of the US public. We then show that rules developed by mental health experts to achieve those objectives enable a RBR that evaluates an LM response's alignment with the objectives similarly to human experts (Pearson's $r=0.841$, $AUC=0.964$). By measuring alignment with objectives that have near unanimous public support, these CoA RBRs provide an approximate measure of alignment between LM behavior and public will.

模型对齐公众意愿心理健康规则奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。