用分步知识驱动优化大模型推理,提升法律等专业领域答题能力
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
- 基于MCTS构建分步监督框架,结合领域知识优化推理过程
- 在法律任务上显著提升准确率,较基线提升12.3个百分点
- 引入反思路径偏好优化,让模型学会从更好视角自我修正
近期,链式思维(CoT)的分步监督在编程与数学等逻辑推理任务中表现优异,借助蒙特卡洛树搜索(MCTS)实现。然而,其在需特定领域知识的任务中的应用仍待探索。本文针对此问题,识别出原始MCTS在该场景下的若干挑战,提出分步领域知识驱动推理优化框架,利用MCTS生成需深度理解、推理与专业知识的问题的分步监督信号。此外,还引入反思路径偏好优化,通过迭代学习更优视角下的自我反思。大量实验验证了方法的有效性,在多个法律领域任务上表现突出。研究还报告了一系列有价值的发现,旨在激发对领域专用大模型与MCTS结合研究的兴趣。
原文摘要 · Abstract (English)
Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search (MCTS). However, its contribution to tasks requiring domain-specific expertise and knowledge remains unexplored. Motivated by the interest, we identify several potential challenges of vanilla MCTS within this context, and propose the framework of Stepwise Domain Knowledge-Driven Reasoning Optimization, employing the MCTS algorithm to develop step-level supervision for problems that require essential comprehension, reasoning, and specialized knowledge. Additionally, we also introduce the Preference Optimization towards Reflection Paths, which iteratively learns self-reflection on the reasoning thoughts from better perspectives. We have conducted extensive experiments to evaluate the advantage of the methodologies. Empirical results demonstrate the effectiveness on various legal-domain problems. We also report a diverse set of valuable findings, hoping to encourage the enthusiasm to the research of domain-specific LLMs and MCTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。