通过增强思维链一致性提升大模型跨领域推理能力
ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- 以思维链步骤间的上下文连贯性替代特定领域知识验证
- 在九个非数学领域平均提升6.5%准确率,显著优于基线
- 适合需要强泛化能力的多领域推理任务
过程奖励模型(PRMs)通过测试时扩展(TTS)显著提升了大语言模型(LLMs)的数学推理能力。然而,多数PRMs在数学领域表现优异,受限于领域特定训练数据稀缺和基于知识的学习模式,在其他领域泛化能力不足。为此,本文将学习目标从验证领域特定知识转向建模领域无关的逻辑流程,聚焦思维链(CoT)步骤间的上下文连贯性,提出一种新型数据标注与训练框架,显著提升模型在多样化领域中的泛化能力。例如,所提出的ContextPRM模型在MMLU-Pro的九个非数学领域(包括法律、历史、哲学)中,通过加权多数投票实现相对于多数投票基线6.5%的平均准确率提升,显著超过VersaPRM的2.2%和其它数学导向PRMs的0.5%提升,展现出在数学与非数学领域的一致优越性能。
原文摘要 · Abstract (English)
Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific training data and knowledge-based learning patterns limits their generalization ability when faced with other domains. To address this limitation, we shift the learning objective from verifying domain-specific knowledge to modeling domain-agnostic logical flow. Centering on contextual coherence between chain-of-thought (CoT) steps, our approach is realized through a novel data annotation and training framework, which enhances the model's generalization capabilities across diverse domains. For instance, our resulting model, ContextPRM, achieves a notable 6.5% average accuracy improvement over the majority voting baseline via weighted majority voting across nine non-mathematical domains in MMLU-Pro, including law, history, and philosophy, significantly surpassing the 2.2% improvement from VersaPRM and 0.5% gains from other mathematics-focused PRMs, demonstrating consistent performance across both mathematical and non-mathematical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。