arXiv:2605.29509cs.CV2026-05

用知识图谱消除文本歧义,实现精准视频生成与编辑

KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing

论文配图:KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
图 1 · 摘自论文原文
  • 构建歧义感知知识图谱,拆解文本为身份、关系等四类语义
  • 在扩散模型关键层注入结构化语义,提升编辑精度与时序一致性
  • 动态调度语义控制,适合高精度文本驱动视频交互场景

近年来,训练自由的视频生成技术发展迅速。然而,在处理复杂文本指令时,现有方法仍面临语义歧义、概念误绑定和跨帧不一致等问题。为此,我们提出KGEdit,一种面向文本到视频扩散模型的结构化语义控制框架。首先,构建歧义感知知识图(AAKG),将输入提示解耦并消歧,转化为身份、关系、属性和负向约束四类结构化语义。随后设计结构化语义注入模块(SSIM),将这些语义信号注入扩散Transformer的关键层,实现细粒度语义控制。此外,引入时序感知语义控制(TASC)模块,根据去噪过程各阶段特性动态调度语义目标,进一步提升语义对齐与时序稳定性。实验表明,KGEdit在编辑精度与时序稳定性方面优于现有方法,同时在文本驱动交互场景中具备更高效率与可控性。

原文摘要 · Abstract (English)

In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still suffer from semantic ambiguity, incorrect concept binding, and cross-frame inconsistency. To address these issues, we propose KGEdit, a structured semantic control framework for text-to-video (T2V) diffusion models. Specifically, we first construct an ambiguity-aware knowledge graph (AAKG) to disentangle and disambiguate the input prompt, converting it into four types of structured semantics: identity, relation, attribute, and negative constraints. We then design a structured semantic injection module (SSIM) to inject these semantic signals into key layers of the diffusion Transformer, enabling fine-grained semantic control. In addition, we introduce a temporal-aware semantic control (TASC) module that dynamically schedules semantic objectives according to the stage-wise characteristics of the denoising process, further improving semantic alignment and temporal consistency. Experiments show that KGEdit outperforms existing methods in editing precision and temporal stability, while offering higher efficiency and controllability in text-driven interaction scenarios.

视频生成知识图谱扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。