在固定算力下,找到视频多模态模型的最佳规模配置。
Inference Compute-Optimal Video Vision Language Models
- 通过大规模实验确定语言模型、帧数、每帧视觉标记数的最优分配。
- 发现任务性能随数据量变化,且最优配置会随之偏移。
- 为资源受限场景提供可落地的模型缩放建议,适合部署优化者。
本研究探讨视频多模态模型中三大关键缩放因子——语言模型规模、帧数、每帧视觉标记数——在推理算力约束下的最优分配。以往工作通常聚焦于模型效率或性能提升,却忽略资源限制。本文在固定推理算力预算下,通过大规模训练扫描与参数化建模,识别出性能-算力的最优边界。实验揭示了任务性能如何依赖于缩放因子及微调数据量,并发现数据量变化会移动最优配置边界。这些发现转化为实际选型建议,指导在资源受限场景中合理选择模型规模。
原文摘要 · Abstract (English)
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the number of visual tokens per frame. While prior works typically focuses on optimizing model efficiency or improving performance without considering resource constraints, we instead identify optimal model configuration under fixed inference compute budgets. We conduct large-scale training sweeps and careful parametric modeling of task performance to identify the inference compute-optimal frontier. Our experiments reveal how task performance depends on scaling factors and finetuning data size, as well as how changes in data size shift the compute-optimal frontier. These findings translate to practical tips for selecting these scaling factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。