不依赖UCB的混合贪心策略更高效地选择生成模型
Mixture-Greedy for Online Generative Model Selection: Is UCB Necessary in Diversity-Aware Multi-Armed Bandits?
- 用混合贪心策略替代传统UCB探索,避免过度试探
- 在FID、Vendi等指标上收敛更快,性能更优
- 适合需要快速筛选生成模型的AI应用
在现代生成式AI中,高效选择多个生成模型至关重要,因使用次优模型采样成本高昂。该问题可建模为多样性感知的多臂老虎机(MAB)任务。以往方法在混合目标中引入上置信界(UCB)探索奖励,但在多个数据集和评估指标下,我们发现UCB项始终延缓收敛并降低采样效率。相反,一种无需显式UCB乐观估计的简单混合贪心策略收敛更快,且在广泛使用的FID和Vendi指标上表现更优,尤其当难以构建紧致置信区间时。理论分析表明:在特定结构条件下,多样性感知目标会通过偏好内部混合体,自然引发探索行为,从而实现所有臂的采样,并对基于多样性的目标获得次线性遗憾。这说明,在多样性感知的多臂老虎机中,如生成模型选择,探索可由目标函数几何结构内生驱动。
原文摘要 · Abstract (English)
Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models is costly. This problem can be viewed as a multi-armed bandit (MAB) task. Under diversity-aware evaluation scores, a non-degenerate mixture of generators can outperform any individual model, distinguishing this MAB setting from classical best-arm identification. Prior approaches incorporate an Upper Confidence Bound (UCB) exploration bonus into the mixture objective. However, across multiple datasets and evaluation metrics, we observe that the UCB term consistently slows convergence and reduces sample efficiency. In contrast, a simple Mixture-Greedy strategy without explicit UCB-type optimism converges faster and achieves even better performance, particularly for widely used metrics such as FID and Vendi where tight confidence bounds are difficult to construct. We provide theoretical insight explaining this behavior: under structural conditions, diversity-aware objectives induce implicit exploration by favoring interior mixtures, leading to sampling of all arms and sublinear regret guarantees for diversity-based objectives. These results suggest that in diversity-aware multi-armed bandits, e.g., for generative model selection, exploration can arise intrinsically from the objective's geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。