arXiv:2508.13104cs.CVcs.RO2025-08ICCV被引 22

用视觉骨骼精准控制复杂人物交互视频生成

Precise Action-to-Video Generation Through Visual Action Prompts

  • 用视觉骨骼作为动作提示,统一表示复杂高自由度动作
  • 在EgoVid、RT-1和DROID数据集上实现高精度动作控制
  • 适合需要精确动作编辑的视频生成研究者

我们提出视觉动作提示(Visual Action Prompts),一种用于复杂高自由度交互动作到视频生成的统一动作表示,同时保持跨域可迁移的视觉动态。现有方法在生成精度与泛化性之间存在权衡:文本、基础动作或粗略掩码虽具泛化性但精度不足,而以智能体为中心的动作信号虽精确却缺乏跨域适应性。为平衡精度与可迁移性,我们提出将动作“渲染”为不依赖特定领域的视觉提示,保留几何精度并支持跨域适配;具体选择视觉骨架因其通用性和易获取性。我们构建了鲁棒流程,从两类交互丰富的数据源——人物交互(HOI)和灵巧机器人操作中提取骨架,实现跨域动作驱动生成模型训练。通过轻量级微调将视觉骨架集成至预训练视频生成模型,实现对复杂交互的精确动作控制,同时保留跨域动态学习能力。在EgoVid、RT-1和DROID上的实验验证了方法的有效性。

原文摘要 · Abstract (English)

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a precision-generality trade-off: existing methods using text, primitive actions, or coarse masks offer generality but lack precision, while agent-centric action signals provide precision at the cost of cross-domain transferability. To balance action precision and dynamic transferability, we propose to "render" actions into precise visual prompts as domain-agnostic representations that preserve both geometric precision and cross-domain adaptability for complex actions; specifically, we choose visual skeletons for their generality and accessibility. We propose robust pipelines to construct skeletons from two interaction-rich data sources - human-object interactions (HOI) and dexterous robotic manipulation - enabling cross-domain training of action-driven generative models. By integrating visual skeletons into pretrained video generation models via lightweight fine-tuning, we enable precise action control of complex interaction while preserving the learning of cross-domain dynamics. Experiments on EgoVid, RT-1 and DROID demonstrate the effectiveness of our proposed approach. Project page: https://zju3dv.github.io/VAP/.

动作生成视频生成视觉骨骼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。