让视频人物声音与形象一致,支持多人场景的个性化生成。
Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation
- 用跨模态绑定和主体锚定提示,统一注入外观与声音特征。
- 在多主体场景下实现音画一致性,效果优于现有方法。
- 适合需要真实人物形象与声音同步的影视创作与虚拟角色开发。
视频合成技术的进步使得真实个体的逼真融合成为可能,推动了对身份感知生成的需求。尽管现有方法可在音视频模型中联合注入外观与声音,但主要集中在单主体场景。多主体间的多模态身份整合仍受限,且多主体场景中视觉与语音身份的精准对齐尚未充分探索。本文提出 Identity-as-Presence,一个统一的联合个性化音视频生成框架。通过自动化数据整理流程构建单/多主体的带身份标签音视频对;采用统一的身份注入机制,通过共享的跨模态身份绑定和主体锚定提示,将外观与声音绑定;多阶段训练策略结合大规模单模态数据与稀缺配对片段,缓解模态不平衡问题。实验表明,在音频质量、视频保真度和音视频一致性上均优于对比方法,尤其在多主体绑定能力上表现更强。
原文摘要 · Abstract (English)
Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerging methods support joint appearance and voice injection in audio-visual models, they primarily focus on single-subject settings. Multimodal identity integration across multiple subjects remains limited, and precise alignment between visual and vocal identities in multi-subject scenarios remains underexplored. We present Identity-as-Presence, a unified framework for joint personalized audio-video generation. An automated data curation pipeline constructs identity-labeled audio-visual pairs for single- and multi-subject scenes. A unified identity injection mechanism then binds paired appearance and voice through shared cross-modal identity binding and subject-anchored captions. A multi-stage training strategy further leverages large-scale unimodal data alongside scarce paired clips to mitigate modality imbalance. Experiments show superior audio quality, video fidelity, and audio-visual consistency, with stronger multi-subject binding than the compared methods. For more details and qualitative results, please refer to our webpage: \href{https://chen-yingjie.github.io/projects/Identity-as-Presence}{Identity-as-Presence}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。