arXiv:2605.04356cs.LGcs.AI2026-05

用在线自然语言反馈提升大模型在模糊领域的对齐效率

Efficiently Aligning Language Models with Online Natural Language Feedback

  • 通过迭代优化代理奖励信号,结合专家少量反馈动态更新
  • 仅需3~50倍更少的人工标注样本,即可恢复80%以上性能
  • 适合资源有限但需高质量对齐的AI训练场景

强化学习结合可验证奖励已在多个领域显著提升语言模型表现。然而,要实现广泛有益的AI部署,需在人类难以直接监督的模糊领域训练具备强能力的模型。本文提出一种方法,在人类专家仅能对少量模型输出提供高质量反馈的条件下,利用在线自然语言反馈实现模型对齐。具体通过迭代优化代理奖励信号,于过拟合前停止,收集新专家反馈并更新代理奖励。代理奖励模型基于大模型通过上下文学习(ICL)和微调构建。实验分别在Qwen3-8B和Haiku 4.5上测试创意写作与对齐研究能力:对于Qwen3-8B,ICL方法以50倍更少样本恢复35%性能,微调方法以20倍样本恢复80%,3倍样本恢复100%;对于Haiku 4.5,ICL方法以30倍样本恢复35%性能,微调方法以10倍样本恢复100%。结果表明,在线自然语言反馈可显著提升专家监督的数据效率。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards has been used to elicit impressive performance from language models in many domains. But, broadly beneficial deployments of AI may require us to train models with strong capabilities in "fuzzy", hard-to-supervise domains. In this paper, we develop methods to align language models in fuzzy domains where human experts are still able to provide high-quality supervision signal, but only for a small number of model outputs, using online natural language feedback. Specifically, we train models by iteratively optimizing against proxy reward signals, stopping at the point of over-optimization, collecting fresh expert supervision, and updating the proxy reward. We construct proxy reward models from language models using in-context learning (ICL) and fine-tuning. We test our methods by eliciting creative writing and alignment research capabilities in Qwen3-8B and Haiku 4.5 respectively. For Qwen3-8B, ICL methods recover up to 35% of performance with 50x fewer expert samples, while fine-tuning methods recover 80% with up to 20x fewer samples and 100% with 3x fewer samples. For Haiku 4.5, ICL methods recover up to 35% of performance with 30x fewer samples, and fine-tuning methods recover 100% with 10x fewer samples. Our results suggest that online natural language feedback can substantially improve the data efficiency of expert supervision.

模型对齐强化学习数据效率在线反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。