用聊天式智能体团队,低成本高效生成叙事幻灯片视频。
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
- 多个智能体分工协作,按脚本、场景、音频流程生成视频。
- 平均成本仅0.103美元,成功率98.4%,计算开销大幅降低。
- 适合内容创作者、教育者快速生成高质量幻灯片视频。
随着人工智能的快速发展,AI生成内容(AIGC)任务显著推动了文本到视频生成领域的发展。然而,传统文本到视频模型通常面临高昂的计算成本。本文提出一种名为视频生成团队(VGTeam)的新颖幻灯片视频生成系统,通过集成大语言模型(LLMs)重构视频创作流程。VGTeam由一组具有通信能力的智能体组成,分别负责脚本撰写、场景生成和音频设计。这些智能体在对话式工作流中协同运作,将用户提供的文本提示转化为连贯的幻灯片风格叙事视频。该系统模拟传统视频制作的阶段性流程,在效率与可扩展性上取得显著提升,同时大幅降低计算开销。平均生成成本仅为0.103美元,成功率达98.4%。该框架在保持高度创意保真度与个性化的同时,为更广泛群体提供高质量内容创作入口,彰显了语言模型在创造性领域的变革潜力,定位为下一代内容创作的先锋系统。
原文摘要 · Abstract (English)
With the rapid advancement of artificial intelligence (AI), the proliferation of AI-generated content (AIGC) tasks has significantly accelerated developments in text-to-video generation. As a result, the field of video production is undergoing a transformative shift. However, conventional text-to-video models are typically constrained by high computational costs. In this study, we propose Video-Generation-Team (VGTeam), a novel slide show video generation system designed to redefine the video creation pipeline through the integration of large language models (LLMs). VGTeam is composed of a suite of communicative agents, each responsible for a distinct aspect of video generation, such as scriptwriting, scene creation, and audio design. These agents operate collaboratively within a chat tower workflow, transforming user-provided textual prompts into coherent, slide-style narrative videos. By emulating the sequential stages of traditional video production, VGTeam achieves remarkable improvements in both efficiency and scalability, while substantially reducing computational overhead. On average, the system generates videos at a cost of only $0.103, with a successful generation rate of 98.4%. Importantly, this framework maintains a high degree of creative fidelity and customization. The implications of VGTeam are far-reaching. It democratizes video production by enabling broader access to high-quality content creation without the need for extensive resources. Furthermore, it highlights the transformative potential of language models in creative domains and positions VGTeam as a pioneering system for next-generation content creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。