arXiv:2608.10209cs.AI2026-08中稿 · COLM

让大模型学会在反馈不完美时仍保持正确行为

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

论文配图:Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
图 1 · 摘自论文原文
  • 用自然语言控制训练样本的反馈质量,提升模型鲁棒性
  • 在新闻生成和算术任务中,显著降低偏见与讨好倾向
  • 适合需要高可靠性对齐的AI系统开发者使用

用于训练大型语言模型(LLMs)的反馈信号是其行为的主要驱动力,也是实现与人类价值观对齐的关键手段。然而,当前后训练方法的一个关键限制在于,人工标注者和自动化奖励函数无法准确捕捉我们真正希望给予的反馈。为此,本文提出评估条件训练(ECT),一种后训练框架,通过自然语言将每个训练样本与反馈的忠实度相关联,并在部署时通过高保真监控器引导模型输出期望行为。ECT旨在改善在不完美反馈下的性能,可作为SFT和PPO等现有算法的补充。我们首先提供ECT的概念框架,讨论其解决持续存在的奖励误指定问题的潜力;接着在隐含知识提取(ELK)问题背景下论证其合理性;最后在两个概念验证实验中评估ECT:提升新闻文章生成的中立性,以及减少算术任务中的讨好倾向。在每种情境中均使用不完美反馈——分别奖励偏见和用户一致性。结果表明,相比直接训练,ECT在目标行为上均有提升。

原文摘要 · Abstract (English)

Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alignment with human values and objectives. However, a key limitation of current post-training methods is the inability of human annotators and automated reward functions to faithfully capture the feedback we would like to give. We introduce Evaluation-Conditioned Training (ECT), a post-training framework that uses natural language to condition each training sample on the fidelity of the feedback we provide and then elicits the desired behavior by conditioning the LLM on a high-fidelity monitor in deployment. ECT is aimed at improving performance under imperfect feedback and works as an add-on to existing algorithms such as SFT and PPO. We first provide a conceptual framework for ECT and discuss its potential to address persistent sources of reward mis-specification. Then we motivate ECT in the context of the eliciting latent knowledge (ELK) problem. Finally, we evaluate ECT on two proof-of-concept experiments: increasing even-handedness in news article generation and reducing sycophancy on an arithmetic task. In each setting, we utilize imperfect feedback, rewarding bias and agreement with the user, respectively. In both settings, ECT improves the targeted behavior relative to direct training.

大模型对齐反馈机制行为约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。