分阶段微调让企业问答模型在保留通用能力下精准掌握私有知识。
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

- 三阶段流程:文档构建任务监督,通用数据回放微调,强化学习优化残差错误。
- 在WnuanBench上,问答准确率从52.76%提升至91.51%,显著提高。
- 残差采样优于全池和随机采样,适合资源受限的企业场景。
企业问答需在不丢失通用能力的前提下掌握私有知识。本文提出Wnuan,一种三阶段流程:从文档构建任务导向监督信号,通过通用数据回放进行监督微调,并利用强化学习优化残差错误。在包含707个问题的WnuanBench测试中,主干32B模型的可接受答案率(AAR)从适应前的52.76%提升至微调后80.06%,强化学习后达91.51%。在匹配的100次更新协议下,残差错误采样相比全池和大小匹配随机采样分别提升3.11和2.97点。源簇自助区间均高于零,同域验证集保持结果排序一致性。通用基准平均下降5.17点,主要集中在指令遵循能力。自动评估集成与领域专家在90.5%的分层样本上达成一致。结果揭示了分阶段企业适配带来的收益与通用能力损失。
原文摘要 · Abstract (English)
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptable-answer rate (AAR) from 52.76% before adaptation to 80.06% after SFT and 91.51% after RL. Under a matched 100-update protocol, residual-error sampling outperforms full-pool and size-matched random sampling by 3.11 and 2.97 points, respectively. Source-cluster bootstrap intervals remain above zero for both contrasts, and a same-domain validation set preserves the ordering. The general-benchmark average decreases by 5.17 points across the route, concentrated in instruction following. The automatic evaluation ensemble agrees with an authoritative domain expert on 90.5% of a stratified Wnuan-Inst response sample. These results characterize both the gains and the general-capability cost of staged enterprise adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。