arXiv:2410.12621cs.CLcs.LG2024-10被引 6

用弱模型引导强模型,提升安全、毒性、法律推理的对齐效果

Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning

  • 用弱监督信号引导强模型生成更符合人类价值的内容
  • 在安全、毒性、法律推理任务中验证了弱到强泛化现象普遍存在
  • 为构建高效对齐方法提供可复现的实验框架和策略建议

随着大语言模型持续进步,确保其与人类价值观对齐变得愈发关键。传统对齐方法严重依赖人工反馈进行微调。然而,当模型性能超越人类理解能力时,仅靠人类判断来评估和对齐这些模型面临巨大挑战。近期工作尝试使用弱监督者从更强模型中提取知识,但现有研究多集中在二分类等简化场景,与真实对齐任务(如安全性)存在显著差异。本文首次将弱到强生成扩展至实际对齐任务,实证揭示在安全、毒性及法律推理三类复杂任务中均存在广泛存在的弱到强泛化现象。同时探索了提升对齐性能的有效策略,并系统分析了各任务中的挑战与潜在解决方案,旨在推动弱到强泛化研究进展。代码已开源。

原文摘要 · Abstract (English)

As large language models (LLMs) continue to advance, ensuring their alignment with human values becomes increasingly critical. Traditional alignment methods heavily rely on human feedback to fine-tune models. With the emergence of superhuman models whose outputs may surpass human understanding, evaluating and aligning these models using human judgments poses significant challenges. To address the challenges, recent works use weak supervisors to elicit knowledge from much stronger models. However, there are important disanalogies between the empirical setup in the existing works and the genuine goal of alignment. We remark that existing works investigate the phenomenon of weak-to-strong generation in analogous setup (i.e., binary classification), rather than practical alignment-relevant tasks (e.g., safety). In this paper, we bridge this gap by extending weak-to-strong generation to the context of practical alignment. We empirically demonstrate the widespread phenomenon of weak-to-strong generation in three complicated alignment tasks: safety, toxicity, and legal reasoning}. Furthermore, we explore efficient strategies for improving alignment performance to enhance the quality of model outcomes. Lastly, we summarize and analyze the challenges and potential solutions in regard to specific alignment tasks, which we hope to catalyze the research progress on the topic of weak-to-strong generalization. Our code is released at https://github.com/yeruimeng/WTS.git.

对齐弱监督安全推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。