系统梳理多模态生成模型的类别与技术,揭示跨模态能力的关键支撑。
A Survey of Generative Categories and Techniques in Multimodal Generative Models
- 按生成模态分类六类,分析自监督、专家混合等核心技术如何实现跨模态生成。
- 提出以忠实性、组合性、鲁棒性为核心的统一评估框架,覆盖多模态基准与人类实验。
- 聚焦生成内容的安全风险,如深度伪造、版权侵权,并提出缓解策略,适合研究者与政策制定者参考。
多模态生成模型(MGMs)已从文本生成拓展至图像、音乐、视频、人体动作和3D物体等多种输出模态,通过在统一架构中融合语言与其他感官模态实现。本综述将主要生成模态分为六类,分析自监督学习(SSL)、专家混合(MoE)、基于人类反馈的强化学习(RLHF)和思维链(CoT)提示等基础技术如何推动跨模态能力的发展。我们梳理关键模型与架构趋势,揭示新兴的跨模态协同效应,提炼可迁移的技术方法并指出未解挑战。基于统一的模型与训练范式分类,提出以忠实性、组合性和鲁棒性为核心的评估框架,综合多模态基准测试与人类研究证据。进一步分析可信度、安全与伦理风险,包括多模态偏见、隐私泄露,以及高保真内容生成在深度伪造、虚假信息传播及音乐与3D资产版权侵犯中的滥用问题,并探讨新兴缓解策略。最后讨论架构趋势、评估协议与治理机制的协同设计,提出缩小能力与安全差距的关键路径,旨在构建更具通用性、可控性与可问责性的多模态生成系统。
原文摘要 · Abstract (English)
Multimodal Generative Models (MGMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with other sensory modalities under unified architectures. This survey categorises six primary generative modalities and examines how foundational techniques, namely Self-Supervised Learning (SSL), Mixture of Experts (MoE), Reinforcement Learning from Human Feedback (RLHF), and Chain-of-Thought (CoT) prompting, enable cross-modal capabilities. We analyze key models, architectural trends, and emergent cross-modal synergies, while highlighting transferable techniques and unresolved challenges. Building on a common taxonomy of models and training recipes, we propose a unified evaluation framework centred on faithfulness, compositionality, and robustness, and synthesise evidence from benchmarks and human studies across modalities. We further analyse trustworthiness, safety, and ethical risks, including multimodal bias, privacy leakage, and the misuse of high-fidelity media generation for deepfakes, disinformation, and copyright infringement in music and 3D assets, together with emerging mitigation strategies. Finally, we discuss how architectural trends, evaluation protocols, and governance mechanisms can be co-designed to close current capability and safety gaps, outlining critical paths toward more general-purpose, controllable, and accountable multimodal generative systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。