arXiv:2502.08234cs.CV2025-02被引 1

提出关键步骤生成新任务,用分步视频简化复杂技能学习。

Learning Human Skill Generators at Key-Step Levels

  • 用多模态大模型提取技能关键步骤描述
  • 通过关键帧图像生成实现步骤间连续性
  • 适合研究具身智能与视频生成的学者

我们致力于在关键步骤层面学习人类技能生成。技能生成极具挑战性,但成功实现将极大促进人类技能学习,并为具身智能提供更丰富经验。尽管当前视频生成模型能合成简单的原子动作,但在处理复杂技能时仍表现不佳,因其涉及多步骤、长时程动作和复杂场景变换,现有自回归方法难以生成完整技能视频。为此,我们提出新任务——关键步骤技能生成(KS-Gen),目标是给定初始状态和技能描述,生成完成该技能的关键步骤视频片段,而非全长视频。为此,我们构建了精心标注的数据集,并定义多种评估指标。针对任务复杂性,我们提出新框架:首先,多模态大语言模型(MLLM)通过检索增强生成关键步骤描述;其次,关键步骤图像生成器(KIG)解决技能视频中关键步骤间的不连续问题;最后,视频生成模型结合描述与关键帧图像,生成具有高时间一致性的关键步骤视频片段。我们对结果进行深入分析,旨在为人类技能生成提供更多洞见。所有模型与数据已开源于https://github.com/MCG-NJU/KS-Gen。

原文摘要 · Abstract (English)

We are committed to learning human skill generators at key-step levels. The generation of skills is a challenging endeavor, but its successful implementation could greatly facilitate human skill learning and provide more experience for embodied intelligence. Although current video generation models can synthesis simple and atomic human operations, they struggle with human skills due to their complex procedure process. Human skills involve multi-step, long-duration actions and complex scene transitions, so the existing naive auto-regressive methods for synthesizing long videos cannot generate human skills. To address this, we propose a novel task, the Key-step Skill Generation (KS-Gen), aimed at reducing the complexity of generating human skill videos. Given the initial state and a skill description, the task is to generate video clips of key steps to complete the skill, rather than a full-length video. To support this task, we introduce a carefully curated dataset and define multiple evaluation metrics to assess performance. Considering the complexity of KS-Gen, we propose a new framework for this task. First, a multimodal large language model (MLLM) generates descriptions for key steps using retrieval argument. Subsequently, we use a Key-step Image Generator (KIG) to address the discontinuity between key steps in skill videos. Finally, a video generation model uses these descriptions and key-step images to generate video clips of the key steps with high temporal consistency. We offer a detailed analysis of the results, hoping to provide more insights on human skill generation. All models and data are available at https://github.com/MCG-NJU/KS-Gen.

技能生成视频生成多模态具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。