arXiv:2601.15224cs.CVcs.CL2026-01ACL被引 8

让视觉语言模型学会判断任务进度,突破静态图像识别局限。

PROGRESSLM: Towards Progress Reasoning in Vision-Language Models

  • 提出两阶段推理框架,结合提示工程与训练优化
  • 仅30亿参数的ProgressLM-3B模型显著提升进度判断准确率
  • 发现模型对视角和演示方式敏感,适合研究长时序推理的学者

任务进度估计需要对长时程动态进行推理,而非仅识别静态视觉内容。尽管现代视觉语言模型(VLMs)擅长描述可见内容,但尚不清楚它们能否从部分观测中推断任务进展程度。为此,我们引入Progress-Bench,一个系统评估VLMs进度推理能力的基准。除基准测试外,我们还探索了受人类启发的两阶段进度推理范式,采用无训练提示和基于精选数据集ProgressLM-45K的训练方法。在14个VLM上的实验表明,多数模型尚未具备任务进度估计能力,对演示模态和视角变化敏感,且难以处理无法回答的情况。虽然无训练提示带来的改进有限且依赖模型,但基于训练的ProgressLM-3B在小规模下仍实现稳定提升,即使其训练任务与评估任务完全不重叠。进一步分析揭示了典型的错误模式,并明确了进度推理成功或失败的条件。

原文摘要 · Abstract (English)

Estimating task progress requires reasoning over long-horizon dynamics rather than recognizing static visual content. While modern Vision-Language Models (VLMs) excel at describing what is visible, it remains unclear whether they can infer how far a task has progressed from partial observations. To this end, we introduce Progress-Bench, a benchmark for systematically evaluating progress reasoning in VLMs. Beyond benchmarking, we further explore a human-inspired two-stage progress reasoning paradigm through both training-free prompting and training-based approach based on curated dataset ProgressLM-45K. Experiments on 14 VLMs show that most models are not yet ready for task progress estimation, exhibiting sensitivity to demonstration modality and viewpoint changes, as well as poor handling of unanswerable cases. While training-free prompting that enforces structured progress reasoning yields limited and model-dependent gains, the training-based ProgressLM-3B achieves consistent improvements even at a small model scale, despite being trained on a task set fully disjoint from the evaluation tasks. Further analyses reveal characteristic error patterns and clarify when and why progress reasoning succeeds or fails. Website: https://progresslm.github.io/ProgressLM/

视觉语言模型进度推理长时序建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。