让文本生成视频准确表达否定语义,无需重训练模型。
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion
- 将否定语义转化为扩散过程中的结构化约束条件
- 在多个否定类型上实现高合规性且保持画面质量
- 适用于预训练模型,适合研究语言与视觉生成的学者
否定是基本的语言操作,但在基于扩散的生成系统中仍缺乏有效建模。本文提出一种形式化处理方式:将语言否定视为扩散动态中语义引导的结构化可行性约束。不依赖启发式规则或重新训练参数,而是将无分类器引导重解释为语义更新方向,并通过投影到由语言结构导出的凸约束集来强制实现否定。该方法统一处理对象缺失、程度性非反转语义、多重否定组合及作用域敏感歧义等现象。本方法无需训练,兼容预训练扩散主干网络,自然扩展至时序视频生成。此外,我们构建了一个结构化的否定导向基准套件,独立识别生成系统中的不同语言失败模式,推动该领域研究。实验表明,该方法在保持视觉保真度和结构一致性的同时,实现稳健的否定合规性,首次建立超越表示层面评估的扩散模型语言否定统一框架。
原文摘要 · Abstract (English)
Negation is a fundamental linguistic operator, yet it remains inadequately modeled in diffusion-based generative systems. In this work, we present a formal treatment of linguistic negation in diffusion-based generative models by modeling it as a structured feasibility constraint on semantic guidance within diffusion dynamics. Rather than introducing heuristics or retraining model parameters, we reinterpret classifier-free guidance as defining a semantic update direction and enforce negation by projecting the update onto a convex constraint set derived from linguistic structure. This novel formulation provides a unified framework for handling diverse negation phenomena, including object absence, graded non-inversion semantics, multi-negation composition, and scope-sensitive disambiguation. Our approach is training-free, compatible with pretrained diffusion backbones, and naturally extends from image generation to temporally evolving video trajectories. In addition, we introduce a structured negation-centric benchmark suite that isolates distinct linguistic failure modes in generative systems, to further research in this area. Experiments demonstrate that our method achieves robust negation compliance while preserving visual fidelity and structural coherence, establishing the first unified formulation of linguistic negation in diffusion-based generative models beyond representation-level evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。