用文本生成60秒电影,图文音三模态联动。
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
- 分五幕结构+图像音频模型联动生成剧情
- 支持1024x768分辨率,帧率15-30FPS,输出专业级画质
- 适合影视创作、教育演示与工业级自动化视频生成
生成式AI已改变多媒体创作方式,实现从文本输入自动生成电影级视频。本文提出一种方法,利用Stable Diffusion生成高保真图像,GPT-2构建叙事结构,并结合gTTS与YouTube音乐的混合音频流水线。采用五幕剧本框架,辅以线性帧插值、电影级后期处理(如锐化)及音视频同步,实现专业级输出。在搭载GPU的Google Colab环境中,基于Python 3.11开发,支持最高1024x768分辨率与15-30 FPS帧率,提供简单与高级双模式Gradio界面。通过CUDA内存管理与错误处理优化,确保系统稳定。实验表明该方法在视觉质量、叙事连贯性与运行效率方面表现优异,推动了文本到视频合成在创意、教育与工业场景中的应用。
原文摘要 · Abstract (English)
Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable Diffusion for high-fidelity image synthesis, GPT-2 for narrative structuring, and a hybrid audio pipeline using gTTS and YouTube-sourced music. It uses a five-scene framework, which is augmented by linear frame interpolation, cinematic post-processing (e.g., sharpening), and audio-video synchronization to provide professional-quality results. It was created in a GPU-accelerated Google Colab environment using Python 3.11. It has a dual-mode Gradio interface (Simple and Advanced), which supports resolutions of up to 1024x768 and frame rates of 15-30 FPS. Optimizations such as CUDA memory management and error handling ensure reliability. The experiments demonstrate outstanding visual quality, narrative coherence, and efficiency, furthering text-to-video synthesis for creative, educational, and industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。