用哲学辩证法让大模型自我反思,提升推理与创新力。
Self-reflecting Large Language Models: A Hegelian Dialectical Approach
- 将黑格尔辩证法形式化为迭代优化过程,生成对立观点并融合。
- 在GSM-8k等数据集上显著超越基线,开放科学创意质量高。
- 适合研究推理增强与跨领域创新的AI学者使用。
本文提出一种基于黑格尔辩证法的自反思框架,将初始命题与生成的对立观点通过迭代融合形成更全面的思想。该方法应用于两类任务:跨学科(数学、物理、经济、哲学)生成新科学思想,以及通过结构化自批判提升大模型的推理能力。研究对比了动态温度退火(从探索到精炼)与恒定温度两种配置,旨在分析温度策略的影响而非推崇其一。为无专家评估思想有效性,引入多智能体多数投票机制(MAMV),由多个大模型独立判断合成结果的合理性和新颖性。实验表明,在数学(GSM-8k, GSM-hard)、符号推理(GSM-Symbolic)及知识密集型任务(MMLU Pro)上均显著优于基线,开放式科学创意也展现出良好质素。
原文摘要 · Abstract (English)
In this paper, we introduce a self-reflection framework for Large Language Models (LLMs) grounded in the Hegelian Dialectic, a philosophical method in which an initial proposition is challenged by a generated opposition, and both are reconciled into a unified, more comprehensive idea. We formalize this process as an iterative operator over the space of consistent theories and apply it to two complementary tasks:(i) generating novel scientific ideas across domains such as mathematics, physics, economics, and philosophy, and (ii)improving reasoning by enabling LLMs to identify and correct their own errors through structured self-critique. We study generation temperature through two configurations (a dynamic annealing schedule that shifts from creative exploration to refinement, and a constant temperature), to examine the effect of fixed versus dynamic temperature rather than advocate either. To evaluate ideas without domain experts, we introduce Multi-Agent Majority Voting (MAMV), in which multiple LLMs independently assess the validity and novelty of each synthesis. Our experiments show significant gains over baselines on mathematical (GSM-8k, GSM-hard), symbolic (GSM-Symbolic), and knowledge-intensive (MMLU Pro) reasoning, with promising qualitative results in open-ended scientific ideation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。