arXiv:2409.08813cs.CL2024-09ICLR被引 20

用弱版大模型做对齐反馈,效果堪比人工且更省资源

Your Weak LLM is Secretly a Strong Teacher for Alignment

  • 用低成本弱模型生成对齐反馈,替代昂贵的人工标注
  • 实验表明弱模型反馈效果可媲美甚至超过人工标注数据
  • 揭示模型规模对反馈质量影响小,适合大规模可持续对齐

大型语言模型(LLM)能力的迅猛发展凸显了对其对齐的迫切需求,以确保模型行为符合人类价值观与意图。现有对齐框架或依赖高成本的人工参与,或需要巨大计算开销。本文探索了一种可行的折中方案:采用资源消耗远低于顶尖模型的弱型大语言模型,实现比纯人工反馈更高程度的自动化。我们系统性地评估并研究了弱模型在生成对齐反馈方面的能力。实证结果表明,弱模型生成的反馈质量可与全人工标注数据相媲美,甚至更优。研究发现模型规模对反馈有效性影响极小,为可扩展、可持续的对齐策略提供了新思路。为进一步理解弱模型反馈下的对齐机制,我们开展了系列定性与定量分析,揭示了人与弱模型反馈在质量上的差异特征。

原文摘要 · Abstract (English)

The burgeoning capabilities of large language models (LLMs) have underscored the need for alignment to ensure these models act in accordance with human values and intentions. Existing alignment frameworks present constraints either in the form of expensive human effort or high computational costs. This paper explores a promising middle ground, where we employ a weak LLM that is significantly less resource-intensive than top-tier models, yet offers more automation than purely human feedback. We present a systematic study to evaluate and understand weak LLM's ability to generate feedback for alignment. Our empirical findings demonstrate that weak LLMs can provide feedback that rivals or even exceeds that of fully human-annotated data. Our study indicates a minimized impact of model size on feedback efficacy, shedding light on a scalable and sustainable alignment strategy. To deepen our understanding of alignment under weak LLM feedback, we conduct a series of qualitative and quantitative analyses, offering novel insights into the quality discrepancies between human feedback vs. weak LLM feedback.

对齐弱模型自动化反馈生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。