发现视觉语言模型与视频生成模型对齐时会丢失细节语义,影响编辑精度。
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing

- 构建可控视频合成数据集TRACE-Edit,分离关系类编辑任务进行诊断
- 四类模型实测显示结构化语义在对齐中严重退化,最高损失达37%
- 为多模态对齐提供新诊断框架,适合视频生成与跨模态研究者
基于流匹配的视频生成模型越来越多依赖前置的视觉语言模型(VLM)来处理复杂指令式视频编辑。当前范式假设连接模块可无缝将VLM的多模态推理能力映射至扩散模型(DiT)原有的文本嵌入空间。然而我们推测,这种对齐机制构成严重语义瓶颈,导致细粒度结构变量退化。验证此假设极具挑战,因端到端评估混淆了对齐失败与生成错误,且自然数据集缺乏解耦标注。为此,我们提出基于视频组合的受控数据处理流程,构建了聚焦关系型编辑的诊断数据集TRACE-Edit。利用该数据集,设计了一套全面诊断协议,分析现有视频编辑模型中元查询与连接器的两种关键设计。对四类代表性模型的系统评估表明,细粒度结构语义在对齐过程中可能严重退化。研究结果推翻了无损语义传递的假设,揭示了VLM-to-DiT对齐是主要瓶颈,并为未来多模态对齐架构提供了新的诊断基础。
原文摘要 · Abstract (English)
Flow matching based video generative models have been increasingly relying on prepended Vision-Language Models (VLMs) to handle complex, instruction-based video editing. The prevailing assumption underlying this paradigm is that a connector module can seamlessly align the VLM's rich multi-modal reasoning with the original text embedding space of DiTs. However, we hypothesize that this alignment acts as a severe semantic bottleneck, degrading fine-grained structural variables. Verifying this is challenging, as end-to-end evaluations conflate alignment failures with generation errors, and natural datasets lack disentangled annotations. To rigorously investigate this, we propose a controlled data processing pipeline based on video composition that results in TRACE-Edit, a diagnostic dataset focusing on relation-based editing. Leveraging this dataset, we propose a comprehensive diagnostic protocol to analyze two important designs of meta-query and connector in the existing video editing models. Systematic evaluation of four representative model cases reveals that fine-grained structural semantics can be severely degraded during alignment. Our findings overturn the assumption of lossless semantic transfer, identifying the VLM-to-DiT alignment as a major bottleneck and providing a new diagnostic foundation for future multi-modal alignment architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。