arXiv:2512.17796cs.CVcs.AI2025-12中稿 · ECCV

让用户用自然语言控制角色在3D场景中自由行动并交互

CustomX: Unified Character, Action, and Scene Customization in Video World Models

  • 基于预训练视频生成器,实现角色与场景的联合可控生成
  • 支持从行走到物体操作的多样化动作,保持长时间视觉连贯性
  • 适合需要高自由度角色交互的虚拟仿真与内容创作

近期世界模型的发展显著提升了交互式环境模拟能力。现有方法主要分为两类:(1) 静态世界生成模型,构建无主动体的3D环境;(2) 可控实体模型,仅允许单一实体在不可控环境中执行有限动作。本文提出CustomX,结合静态世界生成的现实感与结构基础,扩展可控实体模型以支持用户指定的角色执行开放动作。用户可提供3DGS场景和角色,通过自然语言指令让角色完成从基本移动到以物体为中心的交互等多样化行为,并自由探索环境。CustomX将视频生成建模为条件自回归问题,合成时间连贯且保持原始场景与角色视觉保真度的视频片段。基于预训练视频生成器,我们的训练策略显著增强运动动态性,同时保持动作与角色间的泛化能力。评估涵盖视觉质量、角色一致性、动作可控性及长时程连贯性等多个维度。

原文摘要 · Abstract (English)

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2) controllable-entity models, which allow a single entity to perform limited actions in an otherwise uncontrollable environment. In this work, we introduce CustomX, leveraging the realism and structural grounding of static world generation while extending controllable-entity models to support user-specified characters capable of performing open-ended actions. Users can provide a 3DGS scene and a character, then use natural language to direct the character to perform diverse behaviors, ranging from basic locomotion to object-centric interactions, while freely exploring the environment. CustomX synthesizes temporally coherent video clips that preserve visual fidelity with the provided scene and character, formulated as a conditional autoregressive video generation problem. Built upon a pre-trained video generator, our training strategy significantly enhances motion dynamics while maintaining generalization across actions and characters. Our evaluation covers a broad range of aspects, including visual quality, character consistency, action controllability, and long-horizon coherence.

视频生成世界模型自然语言控制角色交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。