arXiv:2605.12480cs.CVcs.AI2026-05被引 3

提出OmniNFT框架,解决音视频联合生成中的模态对齐与同步难题。

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

论文配图:OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
图 1 · 摘自论文原文
  • 按模态分别分配奖励优势,避免多目标奖励冲突
  • 分层梯度手术,防止视频梯度干扰音频生成
  • 关键同步区域动态加权,提升细粒度对齐效果

联合音视频生成近年进展显著,但真实应用需强模态保真、跨模态对齐与细粒度同步。强化学习(RL)虽具潜力,但在多目标、多模态场景下尚未探索。我们分析发现三大瓶颈:(i) 多目标奖励不一致,多模态输出优势在组内不统一;(ii) 多模态梯度失衡,视频分支梯度泄漏至浅层音频生成层;(iii) 统一信用分配,导致细粒度对齐区域探索效率低。这些缺陷使通用RL微调常致次优结果。为此,提出OmniNFT——一种模态感知的在线扩散强化学习框架,包含三项创新:(1) 模态级优势路由,将独立奖励优势定向至对应生成分支;(2) 层级梯度手术,选择性切断浅层音频层的视频梯度,保留跨模态交互层梯度;(3) 区域级损失重加权,引导策略优化聚焦音视频同步与细粒度对齐关键区域。在JavisBench与VBench数据集上,基于LTX-2主干模型的实验表明,OmniNFT显著提升音视频感知质量、跨模态对齐与同步性能。

原文摘要 · Abstract (English)

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) offers a promising paradigm, but its extension to multi-objective and multi-modal joint audio-video generation remains unexplored. Notably, our in-depth analysis first reveals that the primary obstacles to applying RL in this stem from: (i) multi-objective advantages inconsistency, where the advantages of multimodal outputs are not always consistent within a group; (ii) multi-modal gradients imbalance, where video-branch gradients leak into shallow audio layers responsible for intra-modal generation; (iii) uniform credit assignment, where fine-grained cross-modal alignment regions fail to get efficient exploration. These shortcomings suggest that vanilla RL fine-tuning strategy with a single global advantage often leads to suboptimal results. To address these challenges, we propose OmniNFT, a novel modality-aware online diffusion RL framework with three key innovations: (1) Modality-wise advantage routing, which routes independent per-reward advantages to their respective modality generation branches. (2) Layer-wise gradient surgery, which selectively detaches video-branch gradients on shallow audio layers while retaining those for cross-modal interaction layers. (3) Region-wise loss reweighting, which modulates policy optimization toward critical regions related to audio-video synchronization and fine-grained alignment. Extensive experiments on JavisBench and VBench with the LTX-2 backbone demonstrate that OmniNFT achieves comprehensive improvements in audio and video perceptual quality, cross-modal alignment, and audio-video synchronization.

音视频生成强化学习扩散模型多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。