用专家与新手的反馈对比,提升大模型推理能力
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
- 让大模型和小模型分别评价输出,对比生成优化反馈
- 在故事生成、数学推理等任务上最高提升19.6%
- 适合需要高质量推理与内容安全的应用场景
我们提出CLEAR(Contrasting Textual Feedback with Experts and Amateurs for Reasoning),一种新型语言模型推理方法,利用大型(专家)模型与小型(业余)模型的优势。专家与业余模型各自对模型初始输出提供反馈,并相互对比生成优化反馈。该反馈被用于迭代改进CLEAR的输出。实验表明,CLEAR在多个挑战性推理任务中优于现有最佳方法:故事大纲优化(趣味性最高提升19.6%)、受限生成(覆盖率最高提升18.5%)、数学推理(准确率最高提升6.7%)以及毒性缓解(毒性下降最高达22%)。
原文摘要 · Abstract (English)
We introduce CLEAR (Contrasting Textual Feedback with Experts and Amateurs for Reasoning), a novel approach to language model reasoning that leverages the strengths of a larger (expert) model and smaller (amateur) model. The expert and amateur models each provide feedback on a model's initial output and are contrasted with each other into refined feedback. This feedback is subsequently applied to iteratively improve CLEAR's responses. Our experiments demonstrate that CLEAR outperforms state-of-the-art methods in several challenging reasoning tasks, including story outline improvement (up to 19.6% relative increase in interestingness), constrained generation (up to 18.5% increase in coverage), mathematical reasoning (up to 6.7% improvement in accuracy) and mitigation of toxicity (decrease of up to 22% in toxicity).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。