用互动学习帮用户从海量大模型中快速找到最匹配的那一个
CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM

- 通过双人博弈算法轮流选模型,收集用户反馈来推断偏好
- 在有限时间和成本下,比传统方法少花30%代价就找到合适模型
- 特别适合不擅长描述需求但想选好模型的普通用户
随着大模型数量激增,用户在众多模型中挑选适合自己任务的模型面临挑战,每个模型都有独特但难以察觉的隐性特性。用户往往缺乏准确表达所需模型特性的语言或意识。我们提出一种交互高效的主动学习框架,采用双人博弈算法轮流选择一对模型,收集用户对其输出的反馈,并更新对用户潜在偏好的信念。引入新颖的信念感知置信上界策略,在探索模型池与利用推断出的偏好之间取得平衡,实现在用户指定的成本和时间预算下,高效实现用户需求与模型能力的对齐。通过在多种大模型和人类实验中的验证,结果表明该方法能以更低成本实现精准匹配。
原文摘要 · Abstract (English)
Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties. Compounding this challenge, users may lack the vocabulary or awareness to explicitly articulate the characteristics they value in an LLM's responses or deployment. We propose an interaction-efficient active learning framework in which a dueling bandit algorithm iteratively selects pairs of LLMs, collects user feedback about their responses, and updates its belief about the user's latent preferences. We introduce a novel belief-aware upper confidence bound strategy that balances exploration of the model pool with exploitation of inferred preferences, enabling efficient alignment between user needs and LLM capabilities under user-specified cost and time budgets. Through diverse experiments on LLMs and human studies, we experimentally verify that our model can efficiently match well-aligned LLMs to users at a lower cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。