arXiv:2604.07823cs.CVcs.AI2026-04被引 8

让虚拟角色像真人一样实时对话,且长期保持身份一致。

LPM 1.0: Video-based Character Performance Model

  • 用多模态数据训练大模型,实现语音、表情、动作协同控制。
  • 170亿参数模型支持实时生成无限长对话视频,身份稳定不漂移。
  • 专为聊天机器人、直播角色和游戏NPC设计,可实时互动。

表演是通过视觉、声音和时间行为外化意图、情感与个性,使角色鲜活的关键。从视频中学习表演,是替代传统3D制作流程的有前景方向。然而,现有视频模型难以同时实现高表现力、实时推理和长时身份稳定,这一矛盾被称为表演三难困境。对话是最全面的表演场景,角色需同步说话、倾听、反应与表情,并长时间保持身份一致。为此,我们提出LPM 1.0(大型表演模型),专注于单人全双工音视频对话表演。具体而言,我们构建了一个多模态以人为中心的数据集,通过严格筛选、视听配对、表演理解与身份感知的多参考提取;训练一个170亿参数的扩散变换器(基础LPM)以实现高度可控、身份一致的表演;再将其蒸馏为因果流式生成器(在线LPM),实现低延迟、无限长度交互。推理时,给定角色图像与身份感知参考,LPM 1.0能基于用户音频生成倾听视频,基于合成音频生成说话视频,结合文本提示控制动作,全部在实时速度下完成,支持身份稳定、无限时长生成。因此,LPM 1.0可作为对话代理、直播角色和游戏非玩家角色的视觉引擎。为系统评估该场景,我们提出LPM-Bench,首个面向交互式角色表演的基准。LPM 1.0在所有评估维度均达到当前最佳性能,同时保持实时推理。

原文摘要 · Abstract (English)

Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines. However, existing video models struggle to jointly achieve high expressiveness, real-time inference, and long-horizon identity stability, a tension we call the performance trilemma. Conversation is the most comprehensive performance scenario, as characters simultaneously speak, listen, react, and emote while maintaining identity over time. To address this, we present LPM 1.0 (Large Performance Model), focusing on single-person full-duplex audio-visual conversational performance. Concretely, we build a multimodal human-centric dataset through strict filtering, speaking-listening audio-video pairing, performance understanding, and identity-aware multi-reference extraction; train a 17B-parameter Diffusion Transformer (Base LPM) for highly controllable, identity-consistent performance through multimodal conditioning; and distill it into a causal streaming generator (Online LPM) for low-latency, infinite-length interaction. At inference, given a character image with identity-aware references, LPM 1.0 generates listening videos from user audio and speaking videos from synthesized audio, with text prompts for motion control, all at real-time speed with identity-stable, infinite-length generation. LPM 1.0 thus serves as a visual engine for conversational agents, live streaming characters, and game NPCs. To systematically evaluate this setting, we propose LPM-Bench, the first benchmark for interactive character performance. LPM 1.0 achieves state-of-the-art results across all evaluated dimensions while maintaining real-time inference.

视频生成角色表演实时交互扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。