动态分配模型,让复杂任务更省算力还高效
CASTER: Breaking the Cost-Performance Barrier in Multi-Agent Orchestration via Context-Aware Strategy for Task Efficient Routing
- 用语义+结构特征判断任务难易,智能选模型
- 推理成本降72.4%,成功率与强模型持平
- 适合需要高效多智能体协作的工程与科研场景
基于图的多智能体系统(MAS)可支持复杂循环工作流,但静态模型部署导致在简单子任务上浪费大量计算资源。本文提出轻量级路由框架CASTER(上下文感知的任务高效路由策略),通过双信号路由器结合语义嵌入与结构元特征来估计任务难度。训练阶段采用从冷启动到迭代演化的机制,利用自身路由失败的在线策略负反馈自我优化。在软件工程、数据分析、科学发现和网络安全四个领域,基于大模型评分的实验表明,CASTER相比强模型基线最多降低72.4%的推理成本,且成功率相当;同时显著优于启发式路由与FrugalGPT,在所有领域均表现更优。
原文摘要 · Abstract (English)
Graph-based Multi-Agent Systems (MAS) enable complex cyclic workflows but suffer from inefficient static model allocation, where deploying strong models uniformly wastes computation on trivial sub-tasks. We propose CASTER (Context-Aware Strategy for Task Efficient Routing), a lightweight router for dynamic model selection in graph-based MAS. CASTER employs a Dual-Signal Router that combines semantic embeddings with structural meta-features to estimate task difficulty. During training, the router self-optimizes through a Cold Start to Iterative Evolution paradigm, learning from its own routing failures via on-policy negative feedback. Experiments using LLM-as-a-Judge evaluation across Software Engineering, Data Analysis, Scientific Discovery, and Cybersecurity demonstrate that CASTER reduces inference cost by up to 72.4% compared to strong-model baselines while matching their success rates, and consistently outperforms both heuristic routing and FrugalGPT across all domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。