arXiv:2508.13602cs.CV2025-08被引 3

用多智能体协作生成个性化视频日记,支持自我修正与主题化评估

PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

  • 基于多模态大模型构建多智能体协作框架,动态生成高质量内容提示
  • 引入反馈回滚机制,实现多模态内容的迭代自我修正,提升生成质量
  • 提出主题化评测框架ThemeVlogEval,支持标准化评估与公平对比

随着短视频和个性化内容需求的增长,自动化视频日记(Vlog)生成已成为多模态内容创作的关键方向。现有方法多依赖预设脚本,缺乏动态性与个人表达。为此,我们提出PersonaVlog,一种自动化的多模态风格化Vlog生成框架,可根据给定主题和参考图像生成包含视频、背景音乐和内心独白语音的个性化Vlog。该框架基于多模态大语言模型(MLLMs),构建多智能体协作系统,高效生成高质量多模态内容提示,提升创作效率与创意水平。同时,引入反馈与回滚机制,利用MLLMs对生成结果进行评估并提供反馈,实现多模态内容的迭代自我修正。我们还提出了ThemeVlogEval——一个基于主题的自动化评测框架,提供标准化指标与数据集,确保评估公平性。全面实验表明,本框架在多个基线方法上表现显著更优,验证了其在自动化Vlog生成中的有效性与巨大潜力。

原文摘要 · Abstract (English)

With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lacking dynamism and personal expression. Therefore, there is an urgent need for an automated Vlog generation approach that enables effective multimodal collaboration and high personalization. To this end, we propose PersonaVlog, an automated multimodal stylized Vlog generation framework that can produce personalized Vlogs featuring videos, background music, and inner monologue speech based on a given theme and reference image. Specifically, we propose a multi-agent collaboration framework based on Multimodal Large Language Models (MLLMs). This framework efficiently generates high-quality prompts for multimodal content creation based on user input, thereby improving the efficiency and creativity of the process. In addition, we incorporate a feedback and rollback mechanism that leverages MLLMs to evaluate and provide feedback on generated results, thereby enabling iterative self-correction of multimodal content. We also propose ThemeVlogEval, a theme-based automated benchmarking framework that provides standardized metrics and datasets for fair evaluation. Comprehensive experiments demonstrate the significant advantages and potential of our framework over several baselines, highlighting its effectiveness and great potential for generating automated Vlogs.

视频生成多智能体个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。