通过追踪生成轨迹可靠性,实现多模型协同推理。
Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models

- 基于置信度动态识别可靠生成路径,跨模型传递中间结果。
- 在多种推理任务中表现优于单模型,尤其擅长复杂逻辑推理。
- 适合需要多模型协作的高精度生成场景,如智能问答与推理。
掩码扩散语言模型(MDLMs)已成为序列生成的新范式。随着其能力与知识覆盖日益多样,如何融合多模型知识成为关键问题。我们首先研究了MDLM独特的解码动态:成功生成具有稳定答案相关位置的置信度变化,而不可靠轨迹可通过引入其他模型的有前景中间状态进行修正。基于此,提出轨迹迭代集成框架TIE(Trajectory-based Iterative Ensembling),通过跟踪答案相关位置的置信度动态,判断当前哪个模型处于更可靠的生成路径,并选择性地在模型间传递部分去噪序列。由于最优路径随去噪步骤变化,TIE使不同模型在不同阶段贡献互补优势。在多种推理任务上的优异表现及分析表明,TIE为未充分探索的MDLM集成问题提供了一种实用解决方案。
原文摘要 · Abstract (English)
Masked Diffusion Language Models (MDLMs) have emerged as a distinct paradigm for sequence generation. As MDLMs become diverse in capabilities and knowledge coverage, an important question is how to combine their knowledge. Toward this, we first investigate the unique decoding dynamics of MDLMs. We find that successful generations exhibit stable confidence dynamics over answer-relevant positions, while unreliable trajectories can often be corrected by injecting promising intermediate states from other models. Guided by this observation, we propose $\textbf{TIE}$ ($\textbf{T}$rajectory-based $\textbf{I}$terative $\textbf{E}$nsembling), a knowledge fusion framework in which MDLMs iteratively identify reliable decoding trajectories and relay them across models. TIE tracks confidence dynamics over answer-relevant positions to determine which model currently follows a more reliable trajectory and selectively transfers partially denoised sequences across models. As the model on the more promising trajectory often changes across denoising steps, TIE allows different models to contribute complementary strengths at different stages of generation. Strong performance across diverse reasoning tasks, along with our analyses, suggests that TIE offers a practical approach to the underexplored problem of MDLM ensembling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。