用多智能体系统让视频故事角色动态成长,实现沉浸式交互体验。
Facilitating Video Story Interaction with Multi-Agent Collaborative System
- 基于视觉语言模型理解视频剧情,结合检索增强生成与多智能体协作。
- 通过用户提问动态生成角色行为与场景扩展,展现角色成长轨迹。
- 适用于影视互动、个性化叙事创作,适合游戏与教育场景开发。
视频故事交互使观众能够参与并探索叙事内容,以获得个性化体验。然而,现有方法仅限于用户选择、特定设计的叙事,缺乏定制化能力。为此,我们提出一种基于用户意图的交互系统。该系统利用视觉语言模型(VLM)使机器能够理解视频故事,结合检索增强生成(RAG)与多智能体系统(MAS),构建动态演进的角色与场景体验。系统包含三个阶段:1)视频故事处理,借助VLM与先验知识,在三种模态下模拟人类对故事的理解;2)多空间对话,基于用户查询与故事阶段,通过MAS交互生成具有成长性的角色;3)场景定制,根据对话内容扩展并可视化多种故事情节场景。在《哈利·波特》系列上的应用表明,该系统能有效呈现角色社会行为的涌现与成长,显著提升视频故事世界的交互体验。
原文摘要 · Abstract (English)
Video story interaction enables viewers to engage with and explore narrative content for personalized experiences. However, existing methods are limited to user selection, specially designed narratives, and lack customization. To address this, we propose an interactive system based on user intent. Our system uses a Vision Language Model (VLM) to enable machines to understand video stories, combining Retrieval-Augmented Generation (RAG) and a Multi-Agent System (MAS) to create evolving characters and scene experiences. It includes three stages: 1) Video story processing, utilizing VLM and prior knowledge to simulate human understanding of stories across three modalities. 2) Multi-space chat, creating growth-oriented characters through MAS interactions based on user queries and story stages. 3) Scene customization, expanding and visualizing various story scenes mentioned in dialogue. Applied to the Harry Potter series, our study shows the system effectively portrays emergent character social behavior and growth, enhancing the interactive experience in the video story world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。