让大模型选型更省钱又高效,提升任务准确率。
Cost-Aware Model Orchestration for LLM-based Systems
- 用量化性能数据优化大模型选型决策
- 准确率提升0.90%至11.92%,能耗降低54%
- 适合关注成本与效率的AI系统开发者
随着人工智能系统日益复杂,其任务编排常由大型语言模型(LLM)完成,依赖对模型的定性描述进行选择。然而现有描述常不能反映真实性能,导致选型不佳、准确率下降和成本上升。本文通过实证分析揭示了此类编排的局限性,并提出一种考虑性能-成本权衡的感知成本模型选择方法。实验表明,该方法在多个任务上准确率提升0.90%–11.92%,能源效率最高提升54%,且编排选型延迟从4.51秒降至7.2毫秒。
原文摘要 · Abstract (English)
As modern artificial intelligence (AI) systems become more advanced and capable, they can leverage a wide range of tools and models to perform complex tasks. The task of orchestrating these models is increasingly performed by Large Language Models (LLMs) that rely on qualitative descriptions of models for decision-making. However, the descriptions provided to existing LLM-based orchestrators frequently do not reflect true model capabilities and performance characteristics, leading to suboptimal model selection, reduced task accuracy, and increased cost. In this paper, we conduct an empirical analysis of LLM-based orchestration limitations and propose a cost-aware model selection method that accounts for performance-cost trade-offs by incorporating quantitative model performance characteristics within decision-making. Initial experimental results demonstrate that our proposed method increases accuracy by 0.90%-11.92% across various evaluated tasks, achieves up to a 54% energy efficiency improvement, and reduces orchestrator model selection latency from 4.51 s to 7.2 ms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。