arXiv:2512.23562cs.LGcs.AI2025-12中稿 · ed被引 6

首个系统化评估视觉语言模型路由的基准,助力高效多模型选择。

VL-RouterBench: A Benchmark for Vision-Language Model Routing

  • 基于真实推理日志构建质量与成本矩阵,覆盖3大任务组14数据集
  • 包含51.9万样本-模型对,总输入输出token达3449万,评估精度与成本均衡表现
  • 提供可复现工具链,适合研究多模态路由架构与部署优化的开发者

多模型路由已从工程技巧演变为关键基础设施,但现有工作缺乏系统化、可复现的视觉语言模型(VLM)路由评估基准。本文提出VL-RouterBench,用于系统评估VLM路由系统的综合能力。该基准基于原始推理与评分日志,构建样本-模型对的质量与成本矩阵。在规模上,涵盖3个任务组的14个数据集,共30,540个样本,包含15个开源模型和2个API模型,生成519,180个样本-模型对,总输入输出token量为34,494,977。评估协议联合衡量平均准确率、平均成本与吞吐量,并通过归一化后成本与准确率的调和平均构建排名得分,支持不同路由配置与成本预算间的比较。在此基准上,我们评估了10种路由方法及基线,观察到显著的路由收益,但最优现有路由仍与理想Oracle存在明显差距,表明通过更精细的视觉线索与文本结构建模仍有较大优化空间。我们将开源完整数据构建与评估工具链,以促进多模态路由研究中的可比性、可复现性与实际部署。

原文摘要 · Abstract (English)

Multi-model routing has evolved from an engineering technique into essential infrastructure, yet existing work lacks a systematic, reproducible benchmark for evaluating vision-language models (VLMs). We present VL-RouterBench to assess the overall capability of VLM routing systems systematically. The benchmark is grounded in raw inference and scoring logs from VLMs and constructs quality and cost matrices over sample-model pairs. In scale, VL-RouterBench covers 14 datasets across 3 task groups, totaling 30,540 samples, and includes 15 open-source models and 2 API models, yielding 519,180 sample-model pairs and a total input-output token volume of 34,494,977. The evaluation protocol jointly measures average accuracy, average cost, and throughput, and builds a ranking score from the harmonic mean of normalized cost and accuracy to enable comparison across router configurations and cost budgets. On this benchmark, we evaluate 10 routing methods and baselines and observe a significant routability gain, while the best current routers still show a clear gap to the ideal Oracle, indicating considerable room for improvement in router architecture through finer visual cues and modeling of textual structure. We will open-source the complete data construction and evaluation toolchain to promote comparability, reproducibility, and practical deployment in multimodal routing research.

多模态路由评估基准视觉语言模型性能优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。