arXiv:2602.04466cs.CLcs.AI2026-02Conference of the …

微领域自适应预训练能提升大模型查知识,但推理与生成仍存瓶颈。

Is Micro Domain-Adaptive Pre-Training Effective for Real-World Operations? Multi-Step Evaluation Reveals Potential and Bottlenecks

  • 拆解问答为三步:提取事实、推理结论、生成长文
  • 微调后可解决知识提取难题,但推理与生成未提升
  • 需重点加强模型推理能力,适合企业知识应用研究者

将大模型应用于企业实际运维时,需处理特定业务场景中少量专属知识(微领域)。已有研究显示,在小规模文档上进行微领域自适应预训练(mDAPT)效果与大规模领域类似。然而,该研究仅在选择题上评估,其对生成类任务的实际效果尚不明确。本文通过拆解回答过程为三个子任务:(1)从模型自身知识中提取相关事实;(2)基于事实进行推理得出结论;(3)根据结论生成长文本答案,系统评估mDAPT在真实IT技术支持问题上的表现。实验基于企业私有IT产品知识数据集,结果表明:mDAPT有效提升了事实提取能力,但对推理和生成任务无明显改善。进一步分析显示,若事实提取与推理均达标(性能超90%),整体表现即可良好,凸显增强推理能力的必要性。

原文摘要 · Abstract (English)

When applying LLMs to real-world enterprise operations, LLMs need to handle proprietary knowledge in small domains of specific operations ($\textbf{micro domains}$). A previous study shows micro domain-adaptive pre-training ($\textbf{mDAPT}$) with fewer documents is effective, similarly to DAPT in larger domains. However, it evaluates mDAPT only on multiple-choice questions; thus, its effectiveness for generative tasks in real-world operations remains unknown. We aim to reveal the potential and bottlenecks of mDAPT for generative tasks. To this end, we disentangle the answering process into three subtasks and evaluate the performance of each subtask: (1) $\textbf{eliciting}$ facts relevant to questions from an LLM's own knowledge, (2) $\textbf{reasoning}$ over the facts to obtain conclusions, and (3) $\textbf{composing}$ long-form answers based on the conclusions. We verified mDAPT on proprietary IT product knowledge for real-world questions in IT technical support operations. As a result, mDAPT resolved the elicitation task that the base model struggled with but did not resolve other subtasks. This clarifies mDAPT's effectiveness in the knowledge aspect and its bottlenecks in other aspects. Further analysis empirically shows that resolving the elicitation and reasoning tasks ensures sufficient performance (over 90%), emphasizing the need to enhance reasoning capability.

大模型微领域推理能力知识抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。