7B模型实现2K分辨率音视频同步生成,打破传统分离式合成局限。
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

- 采用联合去噪机制,通过门控跨模态注意力耦合音视频流。
- 在2K分辨率下达到与主流开源系统相当的生成质量。
- 开源7B模型与2K精修器,推动音视频生成平民化研究。
现有视频生成模型常忽略音频或分阶段合成,限制视听动态的协同建模。本文提出DreamX-Creator 1.0,一个基于7B参数量的紧凑型原生音视频联合生成系统。该系统以首帧和文本提示为条件,联合去噪音视频专用流。前半网络独立处理,后半通过门控跨模态注意力耦合,其逐令牌与逐头的输出门调控各跨模态注意力头的输出。统一音视频数据系统构建并筛选时序一致的片段,生成结构化多模态标注,并组织成能力导向的数据池。渐进式联合训练包含两个音视频预训练阶段及高质量微调。音视频强化学习进一步使用模态感知的多模态反馈进行后训练,将视频、音频及跨模态反馈分别路由至对应流。针对高分辨率输出,设计自回归1步2K精修管线,将双向多步教师模型转化为自回归多步精修器,并蒸馏为仅需每时间块一次去噪评估的学生模型。整体上,DreamX-Creator 1.0实现了原生、同步的音视频生成,性能媲美顶尖开源系统。通过发布紧凑的7B生成器与2K精修器,旨在推动原生音视频生成的普及,为统一音视频生成建模研究提供可访问基础。
原文摘要 · Abstract (English)
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。