arXiv:2510.27256cs.LGcs.HC2025-10被引 2

根据用户需求自动选模型,快、准、省电兼顾。

ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models

  • 按场景动态选择大模型或小模型,智能分流任务。
  • 80%以上请求由小模型处理,准确率仅降10%以内。
  • 专为视觉语言模型设计,适合边缘计算与高效部署。

视觉语言模型在多种多模态任务中表现优异。然而,用户需求在不同场景下差异明显,可归纳为快速响应、高质量输出和低功耗三类。仅依赖云端大型模型处理所有查询,常导致高延迟和高能耗;而部署于边缘设备的小型模型虽能以低延迟和低功耗处理简单任务,但难以应对复杂需求。为充分发挥大小模型的优势,我们提出 ECVL-ROUTER——首个面向视觉语言模型的场景感知路由框架。该方法引入新的路由策略与评估指标,根据用户需求动态选择适配模型,最大化整体效用。我们构建了一个专用于路由器训练的多模态响应质量数据集,并通过大量实验验证了该方法的有效性。结果表明,我们的方法成功将超过80%的查询路由至小型模型,同时问题求解概率下降不足10%。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel in diverse multimodal tasks. However, user requirements vary across scenarios, which can be categorized into fast response, high-quality output, and low energy consumption. Relying solely on large models deployed in the cloud for all queries often leads to high latency and energy cost, while small models deployed on edge devices are capable of handling simpler tasks with low latency and energy cost. To fully leverage the strengths of both large and small models, we propose ECVL-ROUTER, the first scenario-aware routing framework for VLMs. Our approach introduces a new routing strategy and evaluation metrics that dynamically select the appropriate model for each query based on user requirements, maximizing overall utility. We also construct a multimodal response-quality dataset tailored for router training and validate the approach through extensive experiments. Results show that our approach successfully routes over 80\% of queries to the small model while incurring less than 10\% drop in problem solving probability.

模型路由视觉语言边缘计算高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。