提出一种让自回归模型稳定组合多任务能力的新方法
Compositional Generalization in Autoregressive Models via Logit Composition
- 基于因子化条件概率假设,实现各模型独立控制输出子空间
- 组合后仍保持长度泛化能力,且对输出空间平滑变换不变
- 适合研究模型融合与多技能协同的开发者参考
自回归模型的组合仍是理解大语言模型如何跨任务整合行为或技能的核心挑战。本文提出一种受扩散模型启发的全新、原则性的组合策略。在因子化条件概率假设下,证明组合结果具有投影性:各组件模型可保持对其指定输出子空间的控制,避免相互干扰。该性质在输出空间的光滑重参数化下依然成立,形成特征空间定理。最后,当因子化假设和组件保证在目标长度上一致成立时,组合仍能保持长度泛化行为。这些结果为自回归系统中模型组合与合并的成功提供了原则性理解,并明确了其交互保持稳定的关键条件。
原文摘要 · Abstract (English)
Composing autoregressive models remains a core challenge in understanding how large language models can combine behaviors or skills learned across tasks. We introduce a new and principled composition strategy for autoregressive systems, inspired by composition methods developed for diffusion models. Under a factorized-conditionals assumption, we show that the resulting composition is projective: each component model preserves control over its own designated subspace of the output distribution avoiding interference between models. This property is further preserved under smooth reparameterizations of the output space, yielding a feature-space theorem. Finally, we show that composition preserves length-generalizing behavior when the factorization assumptions and component guarantees hold uniformly at the target length. These results provide a principled understanding of when model composition and merging succeed in autoregressive systems and identify conditions under which their interactions remain stable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。