用视频中的视觉符号直接控制生成,比文字提示更精准。
In-Video Instructions: Visual Signals as Generative Control
- 在视频帧中嵌入文字、箭头等视觉符号作为指令
- 三款主流模型均能准确执行复杂多对象场景的指令
- 适合需要精确空间控制的视频生成任务
大规模视频生成模型近期展现出强大的视觉能力,能够根据当前画面中的逻辑与物理线索预测未来帧。本文探索是否可将此类能力用于可控的图像到视频生成:通过解析帧内嵌入的视觉信号作为指令,提出「视频内指令」(In-Video Instruction)范式。与依赖全局粗略文本提示的方法不同,该范式将用户引导直接编码在视觉域中,如叠加文字、箭头或轨迹,实现对不同物体的明确、空间感知且无歧义的动作指派。在Veo 3.1、Kling 2.5和Wan 2.2三款先进生成器上的大量实验表明,视频模型能可靠理解并执行此类视觉指令,尤其在多对象复杂场景下表现优异。
原文摘要 · Abstract (English)
Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in the current observation. In this work, we investigate whether such capabilities can be harnessed for controllable image-to-video generation by interpreting visual signals embedded within the frames as instructions, a paradigm we term In-Video Instruction. In contrast to prompt-based control, which provides textual descriptions that are inherently global and coarse, In-Video Instruction encodes user guidance directly into the visual domain through elements such as overlaid text, arrows, or trajectories. This enables explicit, spatial-aware, and unambiguous correspondences between visual subjects and their intended actions by assigning distinct instructions to different objects. Extensive experiments on three state-of-the-art generators, including Veo 3.1, Kling 2.5, and Wan 2.2, show that video models can reliably interpret and execute such visually embedded instructions, particularly in complex multi-object scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。