无需训练,通过专家分歧预估文本生成不确定性
EMoE: Training-Free Expert Disagreement for Uncertainty-Aware Text-to-Image Diffusion
- 在MoE扩散模型早期层分离专家路径,利用初始噪声一致性计算潜在表示方差
- 在COCO和CC3M数据集上,提示排序与图文对齐质量相关性优于基线方法
- 可诊断提示风险、模型覆盖范围及多语言生成中的词汇依赖偏差
大型文生图扩散模型通常无法可靠揭示提示生成结果不佳的信号,尤其在训练数据不公开时。本文研究预训练混合专家(MoE)扩散模型中专家分歧是否可作为认知不确定性可靠估计。提出EMoE方法:在早期MoE层分离专家计算路径,保持相同初始噪声,通过首次去噪后潜在表示的方差衡量不确定性。该方法无需额外网络或训练集成模型,即可在完整图像生成前提供不确定性感知提示信号。在COCO和CC3M数据集上,EMoE在提示排序上与图文对齐质量指标的一致性显著优于扩散特定和路由基线方法。进一步应用于多语言提示,发现语言间存在系统性分歧差异与生成质量差异,包括共享词汇效应。这些结果表明,EMoE可作为MoE文生图模型中提示风险、模型覆盖范围及偏见分析的实用诊断工具。
原文摘要 · Abstract (English)
Large text-to-image diffusion models rarely expose reliable signals of when a prompt is likely to produce a poorly aligned generation, especially when training data is undisclosed. We study whether expert disagreement inside pre-trained mixture-of-experts (MoE) diffusion models can serve as a reliable estimate for epistemic uncertainty. We introduce EMoE, a training-free method that separates expert-specific computation paths at an early MoE layer, uses the same initial noise across paths, and measures variance among their latent representations after the first denoising step. This provides an uncertainty-aware prompt signal before full image generation, without auxiliary networks or training diffusion ensembles. On COCO and CC3M, EMoE ranks prompts by text-image alignment quality metrics more consistently than diffusion-specific and router-based baselines. We further apply EMoE to multilingual prompts and find systematic language-dependent differences in disagreement and generation quality, including shared-vocabulary effects. These results position EMoE as a practical diagnostic tool for prompt risk, model coverage, and bias analysis in MoE text-to-image diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。