动态融合专家模型输出,提升推理与编码任务表现
DLLG: Dynamic Logit-Level Gating of LLM Experts

- 在令牌层面动态学习专家模型融合权重
- 无需标签或重训练,跨模型规模表现更优
- 适合需要高效集成多专家模型的场景
利用多个专业化大模型可整合互补优势,但现有方法在适应性与稳定性间权衡:路由过早确定、启发式集成依赖脆弱代理、参数合并引入干扰。我们提出DLLG(动态日志级门控),一种从稀疏响应级监督中学习令牌级专家融合的动态日志级集成框架。轻量级门控模块预测分步融合权重,将轨迹级正确性与生成过程关联,无需令牌级标签或专家重训练。在多样化的推理与代码基准测试中,DLLG在不同模型规模下均持续优于强基线,包括路由、启发式集成和参数合并方法,凸显学习到的日志级融合是一种鲁棒且可扩展的专家集成范式。
原文摘要 · Abstract (English)
Leveraging multiple specialized LLMs can combine complementary strengths, but existing approaches trade adaptability for stability: routing commits prematurely, heuristic ensembling depends on fragile proxies, and parameter merging introduces interference. We propose DLLG (Dynamic Logit-Level Gating), a dynamic logit-level ensembling framework that learns token-level expert fusion from sparse response-level supervision. A lightweight gating module predicts step-wise fusion weights, linking trajectory-level correctness to generation without token-level labels or expert retraining. Across diverse reasoning and code benchmarks, DLLG consistently outperforms strong routing, heuristic ensembling, and parameter-merging baselines across model scales, highlighting learned logit-level fusion as a robust and scalable paradigm for integrating specialized experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。