arXiv:2607.02592cs.CVcs.LG2026-07被引 1

动态分配多教师指导视觉与语言推理,提升多模态模型表现

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

论文配图:H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
图 1 · 摘自论文原文
  • 按解码步骤动态选择最适教师,替代固定教师分配
  • 在11个基准上超越现有方法,显著提升多模态推理准确率
  • 适合需要精准多模态理解的AI系统开发者

近期,基于学生生成轨迹提供监督的在线策略蒸馏(OPD)已成为有效的后训练范式。然而,现有针对多模态推理的OPD方法通常依赖静态教师路由,根据模态或任务类型将每个样本分配给单一教师。这忽略了视觉定位与抽象推理可能在不同解码步骤中占主导地位,导致单一教师无法覆盖完整轨迹。为此,本文提出一种置信度感知的异构多教师在线策略蒸馏框架(H-OPD)。通过验证异构教师在同一推理过程中的互补性,H-OPD 将任务或样本级教师路由替换为沿共享学生轨迹的词元级教师仲裁。H-OPD 利用视觉到语言描述迁移,使纯文本教师可获取关键视觉语义,并采用置信度感知仲裁机制,在每个词元处动态融合视觉-语言教师与纯文本教师。在11个广泛使用的推理基准上的大量实验表明,该方法性能显著优于现有方法。

原文摘要 · Abstract (English)

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods for multimodal reasoning usually rely on a static teacher routing, assigning each sample to a single teacher based on modality or task type. This ignores that visual grounding and abstract reasoning may dominate different decoding steps, making a single teacher insufficient for the full trajectory. To this end, H-OPD is proposed as a confidence-aware heterogeneous multi-teacher OPD framework for multimodal reasoning. By verifying the complementarity of heterogeneous teachers in the same reasoning process, H-OPD replaces task or sample level teacher routing with token-level teacher arbitration along the shared student trajectory. H-OPD employs vision-to-language description transfer to enable text-only teachers to access key visual semantics, and uses a confidence-aware arbitration mechanism to dynamically combine vision-language teacher and text-only teachers at each token. Extensive evaluations over 11 widely-used reasoning benchmarks showcase the superior performance of our method.

多模态推理知识蒸馏教师仲裁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。