用大模型自动生成爆款短视频,效果接近真人内容。
LLMPopcorn: Exploring LLMs as Assistants for Popular Micro-video Generation
- 让大模型充当助手,自动策划并生成流行短视频。
- 优化提示词可进一步提升视频受欢迎程度,最高媲美人工创作。
- 测试显示DeepSeek-V3和R1模型表现最佳,适合内容生成研究者使用。
在抖音、YouTube等平台微视频盛行的背景下,AI生成内容已接近电影级质量。本文首次探索将大语言模型(LLMs)作为助手实现流行微视频的自动化生成(LLMPopcorn)。以爆米花为象征,代表休闲娱乐,契合此类内容常在闲暇时消费的特点。我们实证研究三个问题:(i) 如何有效利用LLM辅助微视频生成?(ii) 基于提示词的增强能否显著提升内容受欢迎程度?(iii) 不同LLM与视频生成器在该任务中的表现如何?结果表明,先进LLM如DeepSeek-V3可生成受欢迎程度媲美人类内容的微视频;提示词优化能进一步提升效果;对比实验表明,DeepSeek-V3和R1在语言模型中表现最优,而LTX-Video和HunyuanVideo在视频生成方面更优。本工作推动了AI辅助微视频创作的发展,并开启新的研究方向。代码已开源:https://github.com/GAIR-Lab/LLMPopcorn。
原文摘要 · Abstract (English)
In an era where micro-videos dominate platforms like TikTok and YouTube, AI-generated content is nearing cinematic quality. The next frontier is using large language models (LLMs) to autonomously create viral micro-videos, a largely untapped potential that could shape the future of AI-driven content creation. To address this gap, this paper presents the first exploration of LLM-assisted popular micro-video generation (LLMPopcorn). We selected popcorn as the icon for this paper because it symbolizes leisure and entertainment, aligning with this study on leveraging LLMs as assistants for generating popular micro-videos that are often consumed during leisure time. Specifically, we empirically study the following research questions: (i) How can LLMs be effectively utilized to assist popular micro-video generation? (ii) To what extent can prompt-based enhancements optimize the LLM-generated content for higher popularity? (iii) How well do various LLMs and video generators perform in the popular micro-video generation task? Exploring these questions, we show that advanced LLMs like DeepSeek-V3 can generate micro-videos with popularity rivaling human content. Prompt enhancement further boosts results, while benchmarking highlights DeepSeek-V3 and R1 for LLMs, and LTX-Video and HunyuanVideo for video generation. This work advances AI-assisted micro-video creation and opens new research directions. The code is publicly available at https://github.com/GAIR-Lab/LLMPopcorn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。