arXiv:2603.27520cs.CV2026-03被引 1

让文本生成视频的属性变化像滑块一样连续可控,还能保持画面连贯。

TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets

  • 在时空视觉令牌空间加偏移量,实现属性调节
  • 无需重训练,可精确控制外观与运动强度
  • 适合需要精细调整视频风格或动作幅度的创作者

我们提出 TokenDial,一种在预训练文本到视频生成模型中实现连续滑块式属性控制的框架。当前生成模型虽能产出高质量整体视频,但在不破坏身份、背景或时间连贯性的前提下,难以精细调节属性变化程度(如效果强度或运动幅度)。TokenDial 的核心观察是:在中间时空视觉补丁令牌空间中添加可加性偏移量,形成语义控制方向,调整偏移量大小即可实现外观与运动动态的连贯、可预测编辑。该方法通过预训练理解信号学习属性特定的令牌偏移量,无需重训练主干网络——外观使用语义方向匹配,运动则采用运动幅度缩放。我们在多种属性和提示上验证了 TokenDial 的有效性,其可控性和生成质量均优于现有最优基线,经量化评估与人类评测支持。

原文摘要 · Abstract (English)

We present TokenDial, a framework for continuous, slider-style attribute control in pretrained text-to-video generation models. While modern generators produce strong holistic videos, they offer limited control over how much an attribute changes (e.g., effect intensity or motion magnitude) without drifting identity, background, or temporal coherence. TokenDial is built on the observation: additive offsets in the intermediate spatiotemporal visual patch-token space form a semantic control direction, where adjusting the offset magnitude yields coherent, predictable edits for both appearance and motion dynamics. We learn attribute-specific token offsets without retraining the backbone, using pretrained understanding signals: semantic direction matching for appearance and motion-magnitude scaling for motion. We demonstrate TokenDial's effectiveness on diverse attributes and prompts, achieving stronger controllability and higher-quality edits than state-of-the-art baselines, supported by extensive quantitative evaluation and human studies.

视频生成连续控制令牌偏移文本到视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。