通过掩码对齐提升视频生成模型对多主体交互的建模能力
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- 在特定层引入掩码轨迹对齐的正则化方法
- 交互相关层注意力集中,显著减少幻觉和漂移
- 适合需要精准交互建模的视频生成研究者
视频扩散模型虽已取得进展,但在建模多实例或主体-客体交互方面仍存挑战。为此,我们构建了包含交互感知描述与多实例掩码轨迹的MATRIX-11K数据集,系统分析视频扩散模型的两种机制:通过视频-文本注意力评估语义定位能力(即名词、动词是否捕捉实例及关系);通过视频-视频注意力评估语义传播能力(即实例绑定是否跨帧持续)。发现二者均集中在少数交互主导层。据此提出MATRIX,一种简单有效的正则化方法,在特定层将视频扩散模型注意力与MATRIX-11K的掩码轨迹对齐,增强定位与传播。进一步提出InterGenEval评估协议。实验表明,MATRIX提升交互保真度与语义对齐,减少漂移与幻觉。消融实验证实设计有效性。代码与权重将公开。
原文摘要 · Abstract (English)
Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, a video dataset with interaction-aware captions and multi-instance mask tracks. Using this dataset, we conduct a systematic analysis that formalizes two perspectives of video DiTs: semantic grounding, via video-to-text attention, which evaluates whether noun and verb tokens capture instances and their relations; and semantic propagation, via video-to-video attention, which assesses whether instance bindings persist across frames. We find both effects concentrate in a small subset of interaction-dominant layers. Motivated by this, we introduce MATRIX, a simple and effective regularization that aligns attention in specific layers of video DiTs with multi-instance mask tracks from the MATRIX-11K dataset, enhancing both grounding and propagation. We further propose InterGenEval, an evaluation protocol for interaction-aware video generation. In experiments, MATRIX improves both interaction fidelity and semantic alignment while reducing drift and hallucination. Extensive ablations validate our design choices. Codes and weights will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。