让大模型学会从错误中反思,提升推理准确性
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
- 引入多视角验证机制,精准优化推理过程
- 利用错误信息提醒模型,避免重复犯错,准确率显著提升
- 在8个数据集上优于所有基线,尤其擅长复杂计算任务
链式思维(Chain-of-Thought, CoT)已成为提升大语言模型性能的关键技术,研究者们致力于通过持续验证与精炼推理输出来改善质量。然而当前范式存在两大问题:(1) 验证方式单一;(2) 忽视推理中的错误信息,每次需从头重构逻辑路径。为此,我们提出错误思维框架(Wrong-of-Thought, WoT),包含两个核心模块:(1) 多视角验证,实现对推理过程与结果的精准修正;(2) 错误信息利用,将错误信息作为警示信号,降低模型重复出错概率。在8个主流数据集和5种大模型上的实验表明,WoT全面超越以往所有基线,尤其在复杂计算任务中表现突出。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) has become a vital technique for enhancing the performance of Large Language Models (LLMs), attracting increasing attention from researchers. One stream of approaches focuses on the iterative enhancement of LLMs by continuously verifying and refining their reasoning outputs for desired quality. Despite its impressive results, this paradigm faces two critical issues: (1) Simple verification methods: The current paradigm relies solely on a single verification method. (2) Wrong Information Ignorance: Traditional paradigms directly ignore wrong information during reasoning and refine the logic paths from scratch each time. To address these challenges, we propose Wrong-of-Thought (WoT), which includes two core modules: (1) Multi-Perspective Verification: A multi-perspective verification method for accurately refining the reasoning process and result, and (2) Wrong Information Utilization: Utilizing wrong information to alert LLMs and reduce the probability of LLMs making same mistakes. Experiments on 8 popular datasets and 5 LLMs demonstrate that WoT surpasses all previous baselines. In addition, WoT exhibits powerful capabilities in difficult computation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。