构建10万+真实手物交互视频数据集,提升机器人模仿学习的泛化能力。
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
- 构建10万+以第一视角记录的手物交互视频数据集
- 微调扩散模型生成逼真交互视频,手部姿态准确率显著提升
- 适合研究机器人模仿学习与视频生成的学者使用
针对现有任务导向手物交互视频生成数据集和模型的关键局限,本文提出TASTE-Rob——首个大规模第一视角手物交互视频数据集,包含100,856段视频。所有视频均与语言指令精确对齐,并从一致视角录制,确保交互清晰。在该数据集上微调视频扩散模型(VDM)后,生成视频具备高度真实性,但手部抓握姿势偶有不一致。为此,我们设计三阶段姿态优化流程,显著改善生成视频中手部姿态的准确性。结合高质量数据集与专用姿态优化框架,实现了更逼真的任务导向手物交互视频生成,显著提升机器人模仿学习的泛化性能。TASTE-Rob数据集及源代码将公开发布于https://taste-rob.github.io。
原文摘要 · Abstract (English)
We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for robotic imitation learning. Current datasets, such as Ego4D, often suffer from inconsistent view perspectives and misaligned interactions, leading to reduced video quality and limiting their applicability for precise imitation learning tasks. Towards this end, we introduce TASTE-Rob -- a pioneering large-scale dataset of 100,856 ego-centric hand-object interaction videos. Each video is meticulously aligned with language instructions and recorded from a consistent camera viewpoint to ensure interaction clarity. By fine-tuning a Video Diffusion Model (VDM) on TASTE-Rob, we achieve realistic object interactions, though we observed occasional inconsistencies in hand grasping postures. To enhance realism, we introduce a three-stage pose-refinement pipeline that improves hand posture accuracy in generated videos. Our curated dataset, coupled with the specialized pose-refinement framework, provides notable performance gains in generating high-quality, task-oriented hand-object interaction videos, resulting in achieving superior generalizable robotic manipulation. The TASTE-Rob dataset is publicly available to foster further advancements in the field, TASTE-Rob dataset and source code will be made publicly available on our website https://taste-rob.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。