用视频构建动态角色画像,让虚拟角色更真实互动
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
- 用视频帧自适应采样,结合动态与静态画像建模角色
- 基于60万条对话数据,生成响应更自然的虚拟角色
- 适合研究虚拟角色交互、多模态对话系统的开发者
角色扮演智能体(RPAs)因其模拟沉浸式互动角色的能力而受到关注。现有方法多依赖静态角色设定,忽视了人类固有的动态感知能力。为此,我们提出将视频模态引入RPAs,构建动态角色画像。为此,我们创建了大规模高质量数据集Role-playing-Video60k,包含60,000个视频和700,000条对应对话。基于此,我们开发了一个综合性框架,融合自适应时间采样机制,结合动态与静态角色画像表示。动态画像通过自适应采样视频帧并按时间顺序输入大语言模型生成;静态画像则由训练视频中的角色对话(用于微调)和输入视频的摘要上下文(用于推理)构成。两者联合建模显著提升了生成响应的质量。此外,我们设计了一套涵盖八个评估指标的稳健评测体系。实验结果验证了该框架的有效性,凸显了动态角色画像在构建高质量角色扮演智能体中的关键作用。
原文摘要 · Abstract (English)
Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic perceptual abilities inherent to humans. To bridge this gap, we introduce the concept of dynamic role profiles by incorporating video modality into RPAs. To support this, we construct Role-playing-Video60k, a large-scale, high-quality dataset comprising 60k videos and 700k corresponding dialogues. Based on this dataset, we develop a comprehensive RPA framework that combines adaptive temporal sampling with both dynamic and static role profile representations. Specifically, the dynamic profile is created by adaptively sampling video frames and feeding them to the LLM in temporal order, while the static profile consists of (1) character dialogues from training videos during fine-tuning, and (2) a summary context from the input video during inference. This joint integration enables RPAs to generate greater responses. Furthermore, we propose a robust evaluation method covering eight metrics. Experimental results demonstrate the effectiveness of our framework, highlighting the importance of dynamic role profiles in developing RPAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。