arXiv:2605.25549cs.CLcs.AI2026-05被引 1

用双专家对话法生成更自然的高质量推理数据。

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

  • 让领域专家与知识工程师协作,通过结构化对话挖掘隐性推理过程。
  • 实验显示新方法生成的推理链自然度得分高达4.80,远超独立写作的1.30。
  • 适合需要深度推理能力的模型训练,尤其在虚构叙事领域效果显著。

高质量专家级思维链(CoT)数据是大语言模型后训练的核心瓶颈。现有方法各有局限:众包标注缺乏深层推理;专家独立写作受'专家盲点'影响,跳过自认为显然的步骤;而强化学习仅生成偏好信号而非完整推理链。本文提出BC协议——一种结构化的双专家对话方法,将领域专家(结晶智力)与知识工程师(流体智力)配对,系统外化专家的隐性判断为自然语言推理链。引入参与者资质模型,定义六维影响因素,并提出原创概念'校准无知'。提出'选择优于规制'原则:在隐性知识挖掘中,资源投入于人选比流程设计回报更高。在叙事虚构领域控制实验中,对比双人对话组(A,n=20)与专家独立写作组(B,n=20),由GPT-4o、Claude Opus 4.5和Gemini 2.5 Pro三模型进行盲评,共600次评分。结果表明,BC协议在推理过程自然度上显著胜出(均值4.80 vs. 1.30,p=2.4×10⁻⁸,Cliff's δ=1.0)。

原文摘要 · Abstract (English)

High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production methods each have structural limitations: crowdsourced annotation lacks deep reasoning paths; expert solo writing is constrained by the "expert blind spot" -- experts structurally skip reasoning steps they consider obvious; RLHF only produces preference signals rather than reasoning chains. This paper proposes the BC Protocol -- a structured dual-expert elicitation method for LLM post-training data production. The method carefully pairs a domain expert (crystallized intelligence) with a knowledge engineer (fluid intelligence), systematically externalizing the expert's implicit judgments as natural language reasoning chains. We introduce the Participant Aptitude Model, which defines six participant characteristic dimensions that affect elicitation quality. "Calibrated Ignorance" is an original concept proposed in this paper. We further propose "Selection-over-Prescription" as a methodological principle: for implicit knowledge elicitation tasks, investing quality-control resources in personnel selection yields a higher return than investing the same resources in process design. In a controlled experiment in the narrative fiction domain, we directly compared CoT produced by BC Protocol dual dialogue (Group A, (n=20)) against CoT written independently by the same domain expert (Group B, (n=20)). Three cross-vendor judge models -- GPT-4o, Claude Opus 4.5, and Gemini 2.5 Pro -- conducted blind evaluation across five dimensions (600 ratings total). Results show that the BC Protocol achieves an overwhelming advantage in "naturalness of reasoning process" (Group A mean 4.80 vs. Group B mean 1.30, (p=2.4\times10^{-8}), Cliff's (δ=1.0)).

思维链双专家数据生成推理质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。