提出按像素级因子动态分配专家,提升离散扩散模型的组合推理能力
From Global to Factor-Wise Expert Composition in Discrete Diffusion Models

- 将生成样本分解为小因子,按需路由给最匹配的专家
- 在ARC-AGI上表现优于全局加权,逻辑一致性和空间解耦更优
- 适合需要精细结构控制的生成任务,如复杂推理与视觉推理
离散扩散模型为解决复杂推理任务提供了强大框架,尤其通过组合生成,将多个预训练专家结合以超越各自训练数据范围。近期理论修正引入时间依赖的混合权重,使组合扩散动态更贴近目标分布。然而,这些方法本质局限在于按样本整体处理,忽略不同专家可能存在的空间或功能专长。本文提出FactorDiff——一种因子级组合框架。我们假设样本可进一步分解为更小因子,并设计一种采样过程,动态将每个因子路由至最相关的专家。我们在空间/像素级组合上实现该框架,并在ARC-AGI基准上验证,结果表明简单因子专属路由在需逻辑一致性与空间解耦的任务中持续优于复杂全局标量加权方案。
原文摘要 · Abstract (English)
Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithically and ignoring the potential spatial or functional specializations of different experts. In this work, we address this limitation by proposing FactorDiff - a factor-wise composition framework for diffusion models. We posit that samples can be further decomposed into smaller factors, and propose a sampling process that dynamically routes each factor to the most relevant expert. We instantiate this framework with spatial/pixel-level compositions and validate it on the ARC-AGI benchmark, demonstrating that simple factor-specific routing consistently outperforms complex global scalar weighting schemes on tasks that require logical consistency and spatial disentanglement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。