arXiv:2506.02726cs.CLcs.AI2025-06

用检索和思维链增强大模型对齐,提升垂直领域推理准确性和可解释性。

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models

  • 结合外部知识检索与思维链生成偏好数据,强化事实依据和推理过程。
  • 在中医领域测试中,准确率和推理深度显著优于基线模型。
  • 适合需要高可信度推理的医疗、法律等专业场景应用。

大型语言模型在垂直领域面临准确性、领域特定推理和可解释性不足的问题。传统偏好对齐方法如基于人类反馈的强化学习(RLHF)和直接偏好优化(DPO)常忽略底层知识来源与推理逻辑。本文提出RACE-Align(检索增强与思维链增强的偏好对齐)框架,系统构建包含外部知识支持与显式思维链(CoT)推理的二元偏好数据集,并采用DPO算法进行模型对齐。核心创新在于偏好数据构造策略:融合AI驱动的知识检索实现事实锚定,提升知识能力与准确性;同时优化领域特异性思维链,将推理过程本身作为关键偏好维度。通过多阶段AI驱动的精炼流水线,高效生成偏好对。以Qwen3-1.7B为基座模型,在中医领域实验表明,RACE-Align显著优于原始基线模型及仅经监督微调(SFT)训练的模型,在答案准确性、信息丰富度、中医思维模式应用、推理逻辑性与深度、可解释性等多个维度均有提升。结果表明,RACE-Align为增强大模型在复杂垂直领域中的知识应用、推理可靠性与过程透明性提供了有效路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) struggle with accuracy, domain-specific reasoning, and interpretability in vertical domains. Traditional preference alignment methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) often overlook the underlying knowledge sources and reasoning logic. This paper introduces RACE-Align (Retrieval-Augmented and Chain-of-Thought Enhanced Alignment), a novel framework designed to address these limitations. RACE-Align systematically constructs a binary preference dataset incorporating external knowledge support and explicit Chain-of-Thought (CoT) reasoning, then aligns LLMs using the DPO algorithm. The core innovation lies in its preference data construction strategy: it integrates AI-driven retrieval for factual grounding, enhancing knowledgeability and accuracy, and emphasizes the optimization of domain-specific CoT, treating the reasoning process itself as a key preference dimension. A multi-stage, AI-driven refinement pipeline cost-effectively generates these preference pairs. Experimental validation in Traditional Chinese Medicine (TCM) using Qwen3-1.7B as the base model demonstrates that RACE-Align significantly outperforms the original base model and a model fine-tuned only with Supervised Fine-Tuning (SFT). Improvements were observed across multiple dimensions, including answer accuracy, information richness, application of TCM thinking patterns, logicality and depth of reasoning, and interpretability. These findings suggest RACE-Align offers an effective pathway to enhance LLMs' knowledge application, reasoning reliability, and process transparency in complex vertical domains.

大模型对齐思维链知识检索中医AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。