用密集视觉监督提升图像视频统一生成的编辑精度与连贯性
From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

- 以深度与法向图为去噪目标,提供结构化视觉监督
- 在统一框架下实现4.15最高综合得分,局部编辑提升至4.18
- 适合需要精准编辑与时序一致性的多媒体生成任务
统一图像与视频生成需同时遵循多样指令并保持身份、几何与时间结构。但仅依赖语义条件和创作训练无法显式监督局部结构,导致精确编辑困难。为此,本文将深度与表面法向预测作为图像去噪目标,利用这些密集任务在统一生成界面中提供结构化监督。框架解耦语义理解与空间对齐的视觉注入,共享一个跨任务的多模态扩散变换器(MMDiT)主干。通过配对视频数据构建方法(MCA)与渐进式训练策略,将学习到的结构线索关联至时空定位编辑与参考条件生成。单个检查点在对比实验中获得4.15的最高综合得分;加入密集监督后,OpenVE整体得分从3.98提升至4.06,局部添加得分从3.92升至4.18。结果表明:感知导向的密集监督能有效传递结构知识,显著增强下游生成任务中的编辑定位与保真度;但不主张其作为独立密集预测模型具有优势。
原文摘要 · Abstract (English)
Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only training do not explicitly supervise the local structure needed for precise, temporally consistent editing. We therefore formulate depth and surface-normal prediction as image-form denoising targets, using these dense tasks as structured visual supervision within the same creation interface. Our framework decouples semantic interpretation from spatially aligned visual injection while sharing one multimodal diffusion transformer (MMDiT) backbone across all tasks. Mutual Context Attention (MCA), a paired-video data-construction procedure, and a progressive training curriculum then connect the learned structural cues to temporally localized editing and reference-conditioned creation. A single checkpoint obtains the highest overall score in the reported comparison of unified systems (4.15); adding dense supervision improves OpenVE Overall from 3.98 to 4.06 and Local Add from 3.92 to 4.18. These results support a deliberately bounded conclusion: perception-oriented dense supervision transfers useful structural knowledge to downstream creation, especially editing locality and preservation; we do not claim superiority as a standalone dense predictor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。