构建首个面向机器人灵巧手的自拍视角图像编辑基准,实现人类手部向机械手的高效转换。
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

- 基于5个数据集构建2亿+编辑样本,支持26种机械手模型的统一编辑任务
- 提出双赛道评估协议,覆盖仅手部与手-臂完整配置,含专用人体感知指标
- 解决人机视觉差异问题,助力从人类视频中低成本训练机器人灵巧操作能力
灵巧手机器人操作是具身智能的核心,但其发展受限于高成本的具身遥操作数据收集。尽管人类手部自拍视频资源丰富,但人手与机器人数据在外观、关节运动和视角上的显著差异,给联合训练带来挑战。现有通用图像编辑模型虽能力强,却缺乏必要的具身先验知识。本文提出HandEdit,一个大规模、具身感知的图像编辑数据集与基准,专用于将自拍视角中的人类手部和手臂转化为多种灵巧机器人形态。HandEdit包含超过2亿个编辑实例,源自五个不同来源数据集,覆盖26种不同的URDF模型(13种仅手部,13种手-臂组合)。我们建立统一的评估协议,分为仅手部和手-臂两个赛道,支持基于URDF条件的评估。通过多维度指标体系(通用相似性度量、基于VLM的判断、具身感知指标)对11个代表性图像编辑基线进行评估。HandEdit作为图像编辑与机器人学交叉的关键资源,推动具身感知编辑模型发展,使利用海量人类视频数据实现可扩展的灵巧机器人学习成为可能,为更泛化的具身智能铺平道路。
原文摘要 · Abstract (English)
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。