arXiv:2604.02319cs.CL2026-04中稿 · COLM被引 1

用路由器选最优模型,提升开放问题的回答多样性。

No Single Best Model for Diversity: Learning a Router for Sample Diversity

  • 根据提示动态选择最适配的模型生成多样回答。
  • 在NB-WildChat上实现26.3%的多样性覆盖率,优于单模型最佳表现。
  • 方法可跨数据集和提示策略泛化,适合多模型协作场景。

当提示允许多个有效答案时,全面生成这些答案是满足多样化用户需求的关键。本文提出多样性覆盖率指标,衡量预测答案集中每个唯一答案相对于最优答案集的质量得分总和。评估18个大模型发现,无单一模型在各类开放性提示中始终领先;但对每个提示,总存在一个显著更优的模型。基于此,我们设计一个路由器,为每条查询推荐最佳模型。在NB-WildChat上,训练后的路由器多样性覆盖率达26.3%,超越单模型最佳基线(23.8%)。该方法进一步在跨域数据集NB-Curated及不同生成提示策略下展现泛化能力。本工作为拥有多个模型时生成全面答案提供了基础框架。

原文摘要 · Abstract (English)

When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In this paper, we study methods to elicit a comprehensive set of valid responses. To evaluate this, we introduce diversity coverage, a metric that measures the total quality scores assigned to each unique answer in the predicted answer set relative to the best possible answer set with the same number of answers. Using this metric, we evaluate 18 LLMs, finding no single model dominates at generating diverse responses to a wide range of open-ended prompts. Yet, per each prompt, there exists a model that outperforms all other models significantly at generating a diverse answer set. Motivated by this finding, we introduce a router that predicts the best model for each query. On NB-WildChat, our trained router outperforms the single best model baseline (26.3% vs 23.8%). We further show generalization to an out-of-domain dataset (NB-Curated) as well as different answer-generation prompting strategies. Our work lays foundation for studying generating comprehensive answers when we have access to a suite of models.

多样性生成模型路由大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。