用法庭辩论机制提升大模型逻辑推理准确率
Judgment-of-Thought Prompting: A Courtroom-Inspired Framework for Binary Logical Reasoning with Large Language Models
- 设计律师、检察官、法官三角色协作,让模型互相辩论
- 在布尔表达式任务中达98%准确率,显著优于传统方法
- 适合需要高可靠性逻辑判断的应用场景
本文提出一种名为思维审判(Judgment of Thought, JoT)的新颖提示方法,专为二元逻辑推理任务设计。尽管提示工程已有进展,现有方法在处理复杂逻辑推理时仍存在局限。为解决此问题,JoT采用多智能体框架,包含律师、检察官和法官三个专门角色:高层模型担任法官,低层模型分别充当律师和检察官,系统性地展开论辩与评估。在BigBenchHard和Winogrande等基准测试中的实验表明,相比现有提示方法,JoT表现更优,尤其在布尔表达式任务中达到98%的准确率。消融实验证明各角色、迭代优化循环及反馈机制均具关键贡献。因此,JoT显著提升了二元推理任务的准确性、可靠性和一致性,具备实际应用潜力。
原文摘要 · Abstract (English)
This paper proposes a novel prompting approach, Judgment of Thought (JoT), specifically tailored for binary logical reasoning tasks. Despite advances in prompt engineering, existing approaches still face limitations in handling complex logical reasoning tasks. To address these issues, JoT introduces a multi-agent approach with three specialized roles$\unicode{x2010}$$\unicode{x2010}$$\unicode{x2010}$lawyer, prosecutor, and judge$\unicode{x2010}$$\unicode{x2010}$$\unicode{x2010}$where a high-level model acts as the judge, and lower-level models serve as lawyer and prosecutor to systematically debate and evaluate arguments. Experimental evaluations on benchmarks such as BigBenchHard and Winogrande demonstrate JoT's superior performance compared to existing prompting approaches, achieving notable improvements, including 98\% accuracy in Boolean expressions. Also, our ablation studies validate the critical contribution of each role, iterative refinement loops, and feedback mechanisms. Consequently, JoT significantly enhances accuracy, reliability, and consistency in binary reasoning tasks and shows potential for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。