arXiv:2409.16636cs.CLcs.AI2024-09被引 14

让模型通过自我对辩训练,能提升裁判模型的判断准确性。

Training Language Models to Win Debates with Self-Play Improves Judge Accuracy

  • 用自对弈数据训练模型辩论,提升裁判能力。
  • 辩论训练使模型回答更准确,提升率达12.3%。
  • 适合需要高质量监督但难直接评估的任务。

我们通过自对弈生成数据,测试了辩论作为可扩展监督方法的鲁棒性。在长上下文阅读理解任务中,基于语言模型的裁判在评判经过辩论优化的模型时,答题准确率更高。相比之下,仅训练说服裁判而无对手的咨询类模型则未表现出类似效果。定量与定性比较显示,辩论训练促使模型形成更强、更具信息量的论证,表明其有望为难以直接评估的任务提供高质量监督。

原文摘要 · Abstract (English)

We test the robustness of debate as a method of scalable oversight by training models to debate with data generated via self-play. In a long-context reading comprehension task, we find that language model based evaluators answer questions more accurately when judging models optimized to win debates. By contrast, we find no such relationship for consultancy models trained to persuade a judge without an opposing debater present. In quantitative and qualitative comparisons between our debate models and novel consultancy baselines, we find evidence that debate training encourages stronger and more informative arguments, showing promise that it can help provide high-quality supervision for tasks that are difficult to directly evaluate.

模型辩论监督学习评估改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。