让大模型模拟五人辩论,提升逻辑推理能力
Town Hall Debate Prompting: Enhancing Logical Reasoning in LLMs through Multi-Persona Interaction
- 用多角色辩论方式让模型从不同视角思考问题
- 5人辩论+LLM生成人格时在ZebraLogic上表现最佳
- 相较单次思维链,推理准确率提升13%以上
辩论是一种高效的问题解决沟通形式,能汇集多元观点。本文提出城镇大会式辩论提示(THDP),将语言模型拆分为多个角色进行相互辩论以得出结论。通过调整角色数量与人格类型,在ZebraLogic这一注重推理的基准上测试表现,该基准包含选择题和填空题。实验表明,使用5个角色且人格由模型自动生成时性能最优:在GPT-4o上,每单元准确率相比单次思维链基线提升13%;Claude 3.5 Sonnet在谜题准确率上提高9%;难题准确率从10%-15%显著上升。
原文摘要 · Abstract (English)
Debate is a commonly used form of human communication catered towards problem-solving because of its efficiency. Debate fundamentally allows multiple viewpoints to be brought up in problem-solving, and for complex problems, each viewpoint opens a new path for problem-solving. In this work, we apply this concept to LLM decision-making by proposing town hall-style debate prompting (THDP), a prompting method that splices a language model into multiple personas that will debate one another to reach a conclusion. Our experimental pipeline varies both the number of personas and the personality types of each persona to find the optimum town hall size and personality for benchmark performance as measured by ZebraLogic bench, a reasoning-intensive benchmark characterized by both multiple-choice and fill-in-the-blank questions. Our experimental results demonstrate that a town hall size of 5 personas with LLM-determined personality types performs optimally on ZebraLogic, achieving a 13\% improvement over one-shot CoT baselines in per-cell accuracy in GPT-4o, 9% puzzle accuracy increase in Claude 3.5 Sonnet, and an improvement in hard puzzle accuracy from 10-15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。