arXiv:2607.24771cs.AIcs.LG2026-07

提出新方法提升知识注入精度,同时保持模型原有能力不变。

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

论文配图:RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
图 1 · 摘自论文原文
  • 用滚动输出条件对比,增强支持事实的生成词权重
  • 在多个测试中准确率优于现有方法,保留能力接近原始模型
  • 适合需要精准注入知识又不破坏原有性能的场景

知识注入可为预训练多模态大模型注入新事实或领域知识,但直接拟合完整权威答案会导致未更新行为发生漂移。在线蒸馏通过基于模型生成的滚动输出进行训练缓解此问题,但统一的参考条件蒸馏提供粗粒度监督:可能低估参考支持的滚动输出词元,并仅间接监督遗漏的事实。我们提出RoCo-ACE,一种针对知识注入的滚动输出条件在线蒸馏目标。RoCo利用同滚动输出的无参考/有参考似然对比,将额外蒸馏权重分配给参考支持的滚动输出词元;而ACE引入稀疏的参考侧锚定修正,对滚动输出中遗漏的权威锚点进行直接修正,无需完整答案模仿。在三种知识注入设置、六个保留性评估基准、多个基线模型及多种基础模型上,RoCo-ACE 在所有对比方法中实现了最高的注入知识准确率,同时评估保留性接近基础模型。

原文摘要 · Abstract (English)

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online distillation mitigates this drift by training on model-generated rollouts, yet uniform reference-conditioned distillation provides coarse supervision: it can under-emphasize reference-supported rollout tokens and supervise omitted facts only indirectly. We introduce RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection. RoCo uses same-rollout reference-free/reference-conditioned likelihood contrast to reallocate additional distillation weight to reference-supported rollout tokens, while ACE adds sparse reference-side anchored correction for authoritative anchors omitted from the rollout without full-answer imitation. Across three knowledge-injection settings, six retention benchmarks, multiple baselines, and multiple base models, RoCo-ACE achieves the best injected-knowledge accuracy among compared methods while keeping evaluated retention close to the base model.

知识注入在线蒸馏多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。