让对话模型学会看视频、思考再说话,提升角色扮演沉浸感
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

- 分三步:看画面→思考逻辑→生成回答,模仿人类反应流程
- 在视觉氛围一致性上比纯文本模型高32%,角色真实度显著提升
- 无需微调即可跨任务使用,适合做VR游戏和互动叙事的开发者
基于文本的角色扮演模型虽能模仿角色风格,但常忽视场景氛围与情绪变化,这对虚拟现实游戏和交互叙事等沉浸式应用至关重要。我们研究视频驱动的角色对话,提出EBM-RL(Eye--Brain--Mouth Reinforcement Learning)框架,采用解耦的GRPO结构,将观察(<perception>)、推理(<think>)与回应生成(<answer>)分离,模拟人类‘看-想-说’过程,使对话基于视觉感知进行推理与生成。为优化该流程,EBM-RL融合场景-文本对齐、感知-认知效用、回答忠实度与格式一致性四类互补奖励。大量实验表明,EBM-RL在自建沉浸式角色扮演基准上显著优于纯文本基线及更大规模视觉语言模型,视觉氛围一致性与角色真实性均大幅提升。此外,该模型在无额外微调下展现出强大的零样本迁移能力,适用于跨域VideoQA任务。我们还开源了首个视频驱动角色对话数据集。
原文摘要 · Abstract (English)
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as VR games and interactive narratives. We study video-grounded role-playing dialogue and introduce EBM-RL (Eye--Brain--Mouth Reinforcement Learning), a decoupled GRPO-based framework that separates observation (<perception>), reasoning (<think>), and utterance generation (<answer>). This design mimics the human See-Think-Speak process, enabling the model to ground dialogue in visual perception before reasoning and response generation. To optimize this See-Think-Speak process, EBM-RL integrates complementary rewards for scene--text alignment, perceptual--cognitive utility, answer faithfulness, and format consistency. Extensive experiments show that EBM-RL substantially outperforms text-only role-playing baselines and larger-scale vision-language models on our immersive role-playing benchmark, improving both visual-atmosphere consistency and character authenticity. Moreover, EBM-RL demonstrates strong zero-shot transfer to out-of-domain VideoQA benchmarks without additional fine-tuning. We also release an open-source dataset for video-grounded role-playing dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。