arXiv:2606.06178cs.LGcs.AI2026-06

让大模型自动匹配用户成本与性能偏好,高效个性化路由。

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

论文配图:Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning
图 1 · 摘自论文原文
  • 用元学习从少量交互中捕捉用户隐式偏好
  • 在分布内/外任务上均优于现有方法,支持多模型路由
  • 适合需要灵活调优大模型成本的开发者和应用方

大型语言模型(LLMs)在性能与成本间存在权衡,更强大的模型带来更高开销。模型路由旨在通过将请求分配给最合适的模型来降低开销并保持性能。然而,现有方法难以适应不同用户的成本-性能偏好。为此,我们提出一种新的感知式模型路由范式,实现个性化、以用户为中心的成本-性能优化,能够通过极少交互高效学习用户的隐式偏好。为应对用户需求异构性,我们将偏好建模为上下文无关的多任务问题,并提出MetaRouter——一个面向偏好感知的元学习框架。实验表明,MetaRouter在分布内与分布外任务上均显著优于强基线,具备高效率的学习能力、对可路由模型变化的鲁棒性以及多模型路由的可扩展性。

原文摘要 · Abstract (English)

Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to mitigate expenses while maintaining performance by sending queries to the most suitable model. However, existing methods cannot perform well for different user cost-performance preferences. To address this gap, we introduce a novel perceptive LLM routing paradigm for personalized and user-centric cost-performance optimization, which efficiently learns users' implicit preferences through little interaction. To handle the challenge of heterogeneous user needs, we formulate preference profiles as a set of distinct tasks in contextual bandit and propose MetaRouter, a meta-learning framework designed for preference-aware LLM routing. Experimental results show that MetaRouter outperforms strong baselines on both in-distribution and out-of-distribution tasks. Furthermore, it exhibits high efficiency in learning user preferences, robustness to changes in the routable LLMs, and scalability to multi-model routing.

大模型路由元学习个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。