arXiv:2608.16513cs.CVcs.AI2026-08

用大语言模型实时纠错,让文字生成视频更准更连贯。

MLLM-Guided Semantic Correction for Text-to-Video Generation

论文配图:MLLM-Guided Semantic Correction for Text-to-Video Generation
图 1 · 摘自论文原文
  • 在生成过程中插入语言模型反馈,动态修正语义偏差。
  • 无需训练,可提升视频语义对齐与时间一致性。
  • 适合需要高精度视频生成的科研与创作场景。

扩散模型与Transformer架构的进步推动了文本到视频生成的发展,但模型常出现物体缺失、属性错误或动作不匹配等语义错误。尽管已有方法在采样前或采样后进行优化,但如何在生成过程中检测并纠正语义偏差仍研究不足。本文提出一种无需训练、可解释的中段生成修正框架,将多模态大语言模型(MLLM)反馈直接融入扩散采样循环。通过在视频合成过程中注入语义评估信号,实现扩散轨迹的修正,使模型能通过持续自省优化生成内容。我们设计两个核心模块:语义评估监督器用于生成中间预览帧以进行语义评估与偏差诊断;语义修改助手通过可控潜空间轨迹干预,在推理阶段修正语义漂移。该方法在不修改模型参数的前提下,提升了语义对齐性、视觉保真度与时间一致性。我们在多个基准上进行了广泛实验验证其有效性。

原文摘要 · Abstract (English)

Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remains underexplored. In this paper, we introduce a training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop. Our framework achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection. We propose two key modules: a Semantic Assessment Supervisor that generates intermediate preview frames for semantic evaluations and deviation diagnostics, and a Semantic Modification Assistant that corrects semantic drift during inference via a controllable latent trajectory intervention. Our method improves semantic alignment, visual fidelity, and temporal consistency without modifying model parameters. We validate the effectiveness of our approach through extensive experiments across multiple benchmarks.

文本生成视频语义纠错扩散模型大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。