视频大模型的性能提升并非均匀分布,部分内容反而因预算增加而变差。
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

- 通过分析每个视频项在不同视觉资源下的响应轨迹,揭示个体差异。
- 40%以上项目在更高预算下准确率下降,最高差距达18.9分。
- 适用于需要精细化评估视频模型的开发者与研究者。
整体规模曲线表明视频大模型随视觉预算增长表现平滑提升或趋于饱和,但我们发现这一观点可能掩盖了个体层面的巨大且相反的变化。通过在受控视觉预算下追踪每个冻结模型-项目对的响应轨迹,我们构建了配置互补性、有害转变和文本覆盖的匹配网格度量。在五个来自三种架构家族的开源视频大模型上,涵盖四组多选题基准、开放式问答与摘要、固定历史对话生成任务,无单一预算能适配所有项目。在四模型匹配多选题网格中,项目级最优潜力跨度为8.8至18.9准确率点,12.5%至25.5%的项目在高预算下反而错误。任务适配的连续指标显示同样互补性:MLVU生成任务的Token-F1最优差距为2.7–3.7分,AVSD当前轮次生成任务为3.8–4.8分,即便平均质量随预算提升。该效应在帧数、空间分辨率、采样策略、时空分配及独立执行的原始视频与缓存流水线中均持续存在,且不受每项速率与成员跟踪协议选择影响。受控采样干预恢复了29.0%的终端退化现象,结构化帧审计识别出若干重复出现的证据路径。我们发布了项目级响应轨迹、协议溯源、衍生标注与可复现分析代码作为审计资产。一种置信度级联方法在保持固定128帧精度的同时,将平均共享帧成本降低31.7%,展示了响应矩阵的一个实际应用场景。
原文摘要 · Abstract (English)
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its response trajectory under controlled visual budgets and derive matched-grid measures of configuration complementarity, harmful transitions, and text overwrite. Across five open Video LLMs from three architecture families, four multiple-choice benchmark splits, open-ended QA and summarization, and fixed-history dialogue generation, no single budget serves all items. On the four-model matched MCQA grid, item-level oracle headroom spans $8.8$--$18.9$ accuracy points and $12.5$--$25.5\%$ of items are correct at a lower budget but wrong at a higher one. Task-appropriate continuous metrics show the same complementarity beyond multiple choice: Token-F1 oracle gaps are $2.7$--$3.7$ score points on MLVU generation and $3.8$--$4.8$ points on AVSD current-turn generation, even when mean quality improves with budget. The effect persists across frame count, spatial resolution, sampling policy, temporal--spatial allocation, and independently executed raw-video and cached pipelines, with per-item rates and membership tracking protocol choices. A controlled sampling intervention recovers $29.0\%$ of terminal regressions, and a structured frame audit identifies several recurring evidence pathways. We release per-item trajectories, protocol provenance, derived annotations, and reproducible analysis code as an auditing artifact. A confidence cascade matches fixed-$128f$ accuracy while reducing average shared frame cost by $31.7\%$, illustrating one operational use of the response matrix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。