提出一个能端到端生成多模态内容的智能体,解决现有模型只能处理单一任务的问题。
A Versatile Multimodal Agent for Multimedia Content Generation
- 基于技能获取理论构建训练数据与代理训练流程
- 采用两阶段关联策略优化生成计划,提升内容一致性
- 适用于视频、音乐等复杂多媒体创作,适合内容生产者使用
随着AIGC技术的发展,生成模型正在革新视频编辑、音乐生成乃至电影制作等领域。然而,当前AIGC模型受限于能力边界,大多仅能在特定场景中作为独立组件使用,难以在真实应用中实现端到端的任务完成。现实中,编辑专家常需处理多种图像与视频输入,生成包含音频、文字等多模态输出的视频内容,这种跨模态整合能力是现有模型难以有效实现的。得益于基于代理的系统兴起,利用AI工具应对复杂内容生成任务成为可能。本文提出MultiMedia-Agent,用于自动化复杂内容创作。该系统包含数据生成管道、内容创作工具库及偏好对齐评估指标。我们引入技能获取理论来建模训练数据筛选与代理训练过程,并设计了包含自相关与模型偏好相关性的两阶段计划优化策略。此外,通过三阶段方法(基础/成功计划微调与偏好优化)利用生成计划训练代理。实验结果表明,所提方法有效,MultiMedia-Agent在生成多媒体内容方面优于新模型。
原文摘要 · Abstract (English)
With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of current AIGC models, most models can only serve as individual components within specific application scenarios and are not capable of completing tasks end-to-end in real-world applications. In real-world applications, editing experts often work with a wide variety of images and video inputs, producing multimodal outputs -- a video typically includes audio, text, and other elements. This level of integration across multiple modalities is something current models are unable to achieve effectively. However, the rise of agent-based systems has made it possible to use AI tools to tackle complex content generation tasks. To deal with the complex scenarios, in this paper, we propose a MultiMedia-Agent designed to automate complex content creation. Our agent system includes a data generation pipeline, a tool library for content creation, and a set of metrics for evaluating preference alignment. Notably, we introduce the skill acquisition theory to model the training data curation and agent training. We designed a two-stage correlation strategy for plan optimization, including self-correlation and model preference correlation. Additionally, we utilized the generated plans to train the MultiMedia-Agent via a three stage approach including base/success plan finetune and preference optimization. The comparison results demonstrate that the our approaches are effective and the MultiMedia-Agent can generate better multimedia content compared to novel models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。