arXiv:2505.16303cs.CL2025-05被引 10

通过能力与知识画像实现大模型高效路由,精准匹配任务需求。

INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling

  • 构建多维能力画像,动态识别最适合的模型
  • 在MMLU-Pro等基准上表现优异,资源利用更高效
  • 适合需要跨模型调度的智能系统开发者

大语言模型(LLM)路由是导航多样化LLM生态的关键技术,旨在根据用户查询领域选择最优性能的专用模型,同时控制计算开销。现有方法在面对大量专业模型时扩展性不足,或难以适应模型范围扩展和能力演化。为此,我们提出InferenceDynamics,一种基于模型能力与知识画像的灵活可扩展多维路由框架。我们在自建数据集RouteMix上进行实验,利用MMLU-Pro、GPQA、BigGenBench和LiveBench等现代基准验证其有效性与泛化能力,证明该框架能准确识别并调用顶尖模型完成任务,在保障高性能的同时实现高效的资源利用。该方法有助于充分释放LLM生态的专精潜力,代码将公开以推动后续研究。

原文摘要 · Abstract (English)

Large Language Model (LLM) routing is a pivotal technique for navigating a diverse landscape of LLMs, aiming to select the best-performing LLMs tailored to the domains of user queries, while managing computational resources. However, current routing approaches often face limitations in scalability when dealing with a large pool of specialized LLMs, or in their adaptability to extending model scope and evolving capability domains. To overcome those challenges, we propose InferenceDynamics, a flexible and scalable multi-dimensional routing framework by modeling the capability and knowledge of models. We operate it on our comprehensive dataset RouteMix, and demonstrate its effectiveness and generalizability in group-level routing using modern benchmarks including MMLU-Pro, GPQA, BigGenBench, and LiveBench, showcasing its ability to identify and leverage top-performing models for given tasks, leading to superior outcomes with efficient resource utilization. The broader adoption of Inference Dynamics can empower users to harness the full specialized potential of the LLM ecosystem, and our code will be made publicly available to encourage further research.

大模型路由能力画像高效调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。