arXiv:2608.18591cs.AIcs.CL2026-08

轻量级模型可预测大模型推理性能,实现文档任务的高效算力分配。

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

  • 用多模态基准预算文档(BudgetDoc)训练轻量级估计器DRB
  • DRB在15组配置中9次优于或持平最大预算基线,成本显著降低
  • 适合追求算力优化的文档理解与大模型部署场景

统一分配大模型推理预算代价高昂且易导致过度计算;尤其在视觉布局复杂的文档任务中。为此,我们提出首个提供显式监督的多模态基准BudgetDoc,涵盖三项文档任务。基于BudgetDoc,训练出约10亿参数的预飞行估计器DRB(SigLIP-2 + Qwen3-0.6B),可预测不同预算水平下的模型性能排序,达到0.753加权F1。在五种前沿模型和三个数据集上动态分配推理预算时,DRB在15组配置中有9组匹配或超越始终使用最大预算的基线,同时大幅降低开销。初步评估显示,DRB具备跨模型选择的泛化潜力。

原文摘要 · Abstract (English)

Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks. Using BudgetDoc, we train DRB (Document-Reasoning Balancer), an approx. 1B-parameter pre-flight estimator (SigLIP-2 + Qwen3-0.6B) that predicts ordinal model performance across budget levels, achieving a 0.753 weighted F1. When dynamically allocating reasoning budgets across five frontier models and three datasets, DRB matches or improves F1 scores compared to always-maximum-budget baselines in 9 of 15 configurations while drastically reducing cost. Finally, preliminary evaluations demonstrate DRB's potential to generalize to cross-model selection.

多模态算力优化推理预算文档理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。