arXiv:2606.23254cs.CVcs.AI2026-06被引 1

让视频文字编辑更精准,还能控制风格和字形。

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

论文配图:SteerVTE: Seamless Video Text Editing with Style and Glyph Control
图 1 · 摘自论文原文
  • 通过风格编码器和双粒度字形编码器控制文本外观与形状。
  • 在100万组数据上训练,视频编辑准确率显著提升。
  • 适合需要精细文字修改的影视、广告等场景使用。

视觉文字编辑旨在精确修改图像和视频中的文字,同时保持风格一致性和视觉真实感。尽管图像领域已有显著进展,视频文字编辑仍处于探索阶段:该任务要求在小范围文字区域实现笔画级精度,进一步加剧了跨帧准确性、时间连贯性与风格保真度的挑战。我们提出SteerVTE,一个统一框架,通过风格与字形控制,引导冻结的视频扩散模型实现精准视频文字编辑。基于冻结的扩散变换器,SteerVTE引入轻量级文本上下文适配器,包含两个互补模块:风格编码器捕捉原始文字的视觉属性,双粒度字形编码器在行与字符层面编码目标文字。为克服视频基础模型固有的弱文本渲染先验,我们提出一种字形感知的空间聚焦损失,并设计三阶段渐进式训练流程,从图像数据逐步扩展到视频数据。为支持大规模训练,我们构建了自动合成流水线,创建了包含一百万组样本的SteerVTE-1M数据集,覆盖多样场景、字体与风格效果。大量实验表明,SteerVTE在文本准确性、风格一致性与时间连贯性方面显著优于现有视频编辑基线方法。

原文摘要 · Abstract (English)

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain, video text editing remains largely unexplored: it is a localized task demanding stroke-level precision within small text regions, which compounds the challenges of cross-frame accuracy, temporal coherence, and stylistic fidelity. We introduce SteerVTE, a unified framework that \underline{\textbf{steer}}s a frozen video diffusion model to perform precise \underline{\textbf{V}}ideo \underline{\textbf{T}}ext \underline{\textbf{E}}diting through style and glyph control. Built on a frozen diffusion transformer, SteerVTE attaches a lightweight text context adapter with two complementary modules: a style encoder capturing the original text's visual attributes, and dual-granularity glyph encoders encoding the target text at both the line and character levels. To overcome the inherently weak text rendering priors of video foundation models, we further propose a glyph-aware spatial-focal loss and a three-stage progressive training curriculum that scales from image to video data. To support large-scale training, we also develop an automatic synthesis pipeline and construct SteerVTE-1M, a dataset of one million triplets spanning diverse scenes, fonts, and stylistic effects. Extensive experiments demonstrate that SteerVTE substantially outperforms existing video editing baselines across text accuracy, style consistency, and temporal coherence.

视频编辑文本生成扩散模型风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。