arXiv:2607.25537cs.CVcs.AI2026-07

通过视觉提示工程提升视频模型推理能力

Visual prompt engineering for video models

论文配图:Visual prompt engineering for video models
图 1 · 摘自论文原文
  • 用图像编辑模型自动优化输入画面以增强视频模型表现
  • 在多个视觉推理任务中,性能显著优于原始图像
  • 适合希望低成本提升视频模型效果的研究者和工程师

在基础模型时代,模型表现很大程度取决于提示质量。随着视频模型逐渐成为视觉任务的基础模型(如视觉推理),我们探究其是否也受益于视觉提示工程——即自动修改任务图像以提升性能。例如,将一个抽象的物理推理场景图转化为逼真版本,仅需调用一次图像编辑模型即可实现。实验表明,视觉提示工程(VIPE)能有效提升多种视频推理任务的表现;对于视频模型而言,其效果甚至优于传统的文本提示工程或测试时扩展方法。这表明,正如文本提示工程系统性地提升语言模型性能,视觉提示工程也可作为简单、高效的手段,激发视频模型更强的视觉推理能力。项目演示视频见 https://visual-prompt-engineering.github.io/。

原文摘要 · Abstract (English)

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.

视频推理提示工程视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。