arXiv:2605.17488cs.CVcs.MM2026-05被引 3

实现音视频身份信息精准绑定,让多人互动生成更自然。

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

论文配图:Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
图 1 · 摘自论文原文
  • 用多模态上下文融合模块增强文本提示,注入视觉与语音特征。
  • 解决语音泄露问题,实现音视频同步且声线、形象保持一致。
  • 适合需要高保真音视频定制的影视创作与虚拟人应用。

联合音视频生成领域因基础模型的兴起而发生根本性变革。尽管取得进展,但在多个交互主体间同时保留视觉身份与声线特征的协同多模态定制仍鲜有研究。为此,我们提出Omni-Customizer,一个端到端框架,旨在精确绑定并无缝融合多模态身份信息。具体地,引入全景上下文融合(OCF)模块,将密集的多模态身份线索融入基础文本提示;设计掩码式语音合成交叉注意力(MTP-CA)机制,有效防止严重的“语音泄露”问题。在架构中,提出语义锚定多模态旋转位置编码(SA-MRoPE),将视觉、音频参考标记及语音合成嵌入锚定至对应语义描述,实现结构化多模态融合与鲁棒的身份绑定。此外,设计综合训练策略:采用交错式音视频调度,快速适应多语言场景而不破坏基础先验;通过从同对到跨对的渐进课程学习,促进高层且稳健的身份特征学习。大量实验表明,Omni-Customizer在双模态定制生成任务上达到当前最优表现,显著优于基线,在视觉身份相似度、声线一致性、音视频精确同步及整体音视频保真度方面均表现优异。

原文摘要 · Abstract (English)

The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides, achieving cohesive multimodal customization for the simultaneous preservation of visual identities and vocal timbres across multiple interacting subjects remains largely underexplored. To bridge this gap, we present Omni-Customizer, an end-to-end framework targeted at the precise binding and seamless fusion of multimodal identity information. Specifically, we introduce an Omni-Context Fusion (OCF) module that effectively enriches the base textual prompt with dense, multimodal identity cues, along with a Masked TTS Cross-Attention (MTP-CA) mechanism explicitly designed to prevent the severe "speech leakage" problem. Within this architecture, we propose Semantic-Anchored Multimodal RoPE (SA-MRoPE) to anchor visual and audio reference tokens, along with TTS embeddings, to their corresponding semantic descriptions, enabling structured multimodal fusion and robust identity binding. Furthermore, we devise a comprehensive training strategy that incorporates interleaved audio-video scheduling to rapidly adapt the audio branch to multilingual scenarios without degrading foundational priors, and a progressive in-pair to cross-pair curriculum to facilitate the learning of high-level and robust identity features. Extensive experiments demonstrate that Omni-Customizer achieves state-of-the-art performance in dual-modal customized generation, excelling across visual identity similarity, timbre consistency, precise audio-video synchronization, and overall video-audio fidelity.

音视频生成多模态定制身份绑定联合生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。