让虚拟人像能自主规划并执行长期任务,实现真正智能交互。
Active Intelligence in Video Avatars via Closed-loop World Modeling
- 构建闭环推理循环,实时验证生成结果以维持状态追踪。
- 在开放场景中完成多步任务,成功率显著优于传统方法。
- 适合研究智能体决策、视频生成与具身认知的学者使用。
当前视频虚拟人生成方法在身份保持和动作对齐方面表现优异,但缺乏真正的自主性,无法通过适应性环境交互实现长期目标。为此,我们提出 L-IVA(长时交互视觉虚拟人)任务与基准,以及 ORCA(在线推理与认知架构),首个实现视频虚拟人主动智能的框架。ORCA 通过两项核心创新实现内部世界模型能力:(1) 闭合回路的 OTAR 循环(观察-思考-行动-反思),通过持续将预测结果与实际生成对比,增强在生成不确定性下的状态追踪鲁棒性;(2) 分层双系统架构:系统 2 执行战略推理并预测状态,系统 1 将抽象计划转化为精确的模型特定动作描述。通过将虚拟人控制建模为部分可观测马尔可夫决策过程,并结合结果验证实现连续信念更新,ORCA 实现了开放域场景下的自主多步任务完成。大量实验表明,ORCA 在任务成功率与行为一致性上显著优于开环及无反思基线,验证了受内部世界模型启发的设计,推动视频虚拟人从被动动画迈向主动、目标导向行为。
原文摘要 · Abstract (English)
Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency, they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and benchmark for evaluating goal-directed planning in stochastic generative environments, and ORCA (Online Reasoning and Cognitive Architecture), the first framework enabling active intelligence in video avatars. ORCA embodies Internal World Model (IWM) capabilities through two key innovations: (1) a closed-loop OTAR cycle (Observe-Think-Act-Reflect) that maintains robust state tracking under generative uncertainty by continuously verifying predicted outcomes against actual generations, and (2) a hierarchical dual-system architecture where System 2 performs strategic reasoning with state prediction while System 1 translates abstract plans into precise, model-specific action captions. By formulating avatar control as a POMDP and implementing continuous belief updating with outcome verification, ORCA enables autonomous multi-step task completion in open-domain scenarios. Extensive experiments demonstrate that ORCA significantly outperforms open-loop and non-reflective baselines in task success rate and behavioral coherence, validating our IWM-inspired design for advancing video avatar intelligence from passive animation to active, goal-oriented behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。