arXiv:2608.26735cs.CL2026-08

让专业模型保持通用能力,避免专精时变傻。

Preserving General Capabilities during Domain Specialization with Uncertainty-Calibrated MOPD

  • 用双温度采样扩大学习轨迹范围,筛选强正向信号。
  • 通过熵校准的教师认可度,保留方向一致的更新令牌。
  • 在角色扮演和医疗领域提升通用能力,效果显著且不牺牲专业性能。

将大语言模型专业化于垂直领域可提升特定任务表现,但常导致推理、编程、指令遵循和创意写作等通用能力下降。本文研究多教师在线策略蒸馏(MOPD)中的这一权衡问题,其中学生模型由领域与通用教师共同指导其自采轨迹。标准MOPD存在两个局限:普通在线采样难以暴露具有大正向教师-学生优势的词元;仅凭优势符号无法判断更新方向是否可靠。为此提出不确定性校准的MOPD:采用双温度采样扩展候选轨迹池,正向优势密度过滤筛选强正向学习信号;中心对数似然(CLL)过滤计算熵校准的教师认可度,按方向-认可一致性概率保留词元更新。在角色扮演与医疗领域专业化实验中,相比标准MOPD,本方法使通用能力平均提升4.73%和10.84%,同时保持垂直领域性能。消融与诊断分析表明,增益并非源于更大回滚预算,且所提轨迹与词元级机制有效解决了预设失效模式。

原文摘要 · Abstract (English)

Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domain--general trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized student is supervised on its own sampled trajectories by domain and general teachers. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, while the advantage sign alone does not establish whether the resulting update direction is reliable. We propose uncertainty-calibrated MOPD to address these limitations. Dual-temperature sampling broadens the candidate trajectory pool, and positive-advantage-density filtering selects trajectories with stronger positive learning signals. Centered log-likelihood (CLL) filtering then computes an entropy-calibrated teacher-endorsement score and probabilistically retains token updates according to direction--endorsement consistency. Experiments on role-playing and medical-domain specialization show that our method improves the general-capability average over standard MOPD by $4.73\%$ and $10.84\%$, respectively, while maintaining vertical-domain performance. Ablations and diagnostic analyses further confirm that the gains do not merely result from a larger rollout budget and that the proposed trajectory- and token-level mechanisms address their intended failure modes.

模型蒸馏通用能力领域专精不确定性校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。