通过反事实敏感性识别关键响应位置,提升语言模型的有选择性强化学习效果。
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

- 基于反事实与改写校准,衡量每个词元的任务相关性。
- 在两种师生设置中分别提升1.92和2.96分,优于现有方法。
- 适合需要精准监督信号的强化学习场景,如对话系统优化。
在线策略蒸馏(OPD)对当前策略采样轨迹中的响应词元进行统一监督,但未区分其监督价值。有选择性的OPD通过估计训练价值非均匀分配监督,然而多数现有标准仅关注优化需求(如不确定性或师生分歧),而忽略任务相关性——即监督是否与输入语义内容紧密关联。为此,本文提出反事实相关性蒸馏(CROP),通过改写校准的反事实敏感性边际来量化任务相关性。针对每个源提示,构建经过验证的原句-改写-反事实三元组,固定学生模型输出,以词元对任务相关性变化的敏感度衡量其重要性,并用对语义保持改写的敏感度进行校准。对比实验表明,CROP比随机或低相关性选择更有效识别高价值监督位置;组件分析证实反事实敏感性和改写校准均具关键作用。在两种师生设置下,CROP相较最强非CROP选择器分别提升1.92和2.96分。结果验证了任务相关性作为补充标准的价值,并确立了CROP为模型内、对比特定的词元级监督分配方法。
原文摘要 · Abstract (English)
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimated training value. Most existing criteria, however, focus primarily on optimization need, such as uncertainty or teacher-student disagreement, while task relevance, namely whether the supervision is tied to the semantic content of the current input, remains less directly characterized as a complementary dimension. To address this gap, we introduce Counterfactual Relevance for On-Policy Distillation (CROP), which operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original-paraphrase-counterfactual triplet, holds the student rollout fixed, and measures each response position by its sensitivity to a task-relevant condition change calibrated by its sensitivity to a meaning-preserving rewrite. Matched selection controls show that CROP identifies more useful supervision positions than random or lowest-relevance selection, while component comparisons confirm the value of both counterfactual sensitivity and paraphrase calibration. Across two teacher-student settings, CROP improves aggregate performance by 1.92 and 2.96 points over the strongest non-CROP selector. These results support task relevance as a complementary criterion for selective OPD and establish CROP as a model-internal, contrast-specific method for allocating token-level supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。