arXiv:2601.03233cs.CV2026-01被引 174

LTX-2可同步生成音视频,质量媲美闭源模型且更省算力。

LTX-2: Efficient Joint Audio-Visual Foundation Model

  • 双流架构:视频流140亿参数,音频流50亿参数,跨模态注意力协同
  • 在多语言提示下生成匹配场景情绪与环境的自然音效,音画同步精准
  • 开源免费,推理速度比闭源模型快,适合影视、游戏等创意应用

近期文本到视频扩散模型虽能生成逼真视频,但缺乏声音——缺失语义、情感和氛围线索。我们提出LTX-2,一个开源的基础模型,可统一生成高质量、时间同步的音视频内容。该模型采用非对称双流Transformer结构,视频流140亿参数,音频流50亿参数,通过双向音视频交叉注意力层、时序位置编码及跨模态AdaLN实现共享时间步条件控制。此设计提升训练与推理效率,并为视频生成分配更多容量。使用多语言文本编码器增强提示理解,引入模态感知无分类器引导(modality-CFG)机制以改善音视频对齐与可控性。除了语音,模型还生成丰富连贯的音频轨道,包含角色动作、环境背景与拟声效果,贴合场景风格与情绪。评估显示,其音视频质量与提示遵循度在开源系统中处于领先地位,性能接近闭源模型,计算成本与推理时间仅为后者的几分之一。所有模型权重与代码均已公开。

原文摘要 · Abstract (English)

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model capable of generating high-quality, temporally synchronized audiovisual content in a unified manner. LTX-2 consists of an asymmetric dual-stream transformer with a 14B-parameter video stream and a 5B-parameter audio stream, coupled through bidirectional audio-video cross-attention layers with temporal positional embeddings and cross-modality AdaLN for shared timestep conditioning. This architecture enables efficient training and inference of a unified audiovisual model while allocating more capacity for video generation than audio generation. We employ a multilingual text encoder for broader prompt understanding and introduce a modality-aware classifier-free guidance (modality-CFG) mechanism for improved audiovisual alignment and controllability. Beyond generating speech, LTX-2 produces rich, coherent audio tracks that follow the characters, environment, style, and emotion of each scene -- complete with natural background and foley elements. In our evaluations, the model achieves state-of-the-art audiovisual quality and prompt adherence among open-source systems, while delivering results comparable to proprietary models at a fraction of their computational cost and inference time. All model weights and code are publicly released.

音视频生成扩散模型多模态开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。