构建全球最大3D人体数据集,支持日常穿搭与动作捕捉。
MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
- 采集4500人多视角日常穿搭与动作,可扩展性强。
- 含64500万帧、9000套服饰、6万段动作序列及丰富标注。
- 适合3D人体建模、动画生成等研究者使用。
当前大语言模型和文生图模型的成功得益于大规模数据集的推动。然而在3D视觉领域,尽管对象中心任务已通过Objaverse和MVImgNet等数据集取得显著进展,人体中心任务却因缺乏类似规模的人体数据集而发展受限。为此,我们提出MVHumanNet++,一个包含4500个不同身份个体的多视角人体动作序列数据集。其核心在于利用多视角捕捉系统收集具有多样化日常着装的人体数据,实现可扩展的数据采集。该数据集包含9000套日常服饰、60000段动作序列和6.45亿帧图像,附带人体掩码、相机参数、2D/3D关键点、SMPL/SMPLX参数及对应文本描述等丰富标注。此外,还新增了处理后的法向量图与深度图,大幅拓展其在人体相关研究中的应用潜力。为验证其价值,我们开展了多项初步实验,展示了其在多种2D/3D视觉任务中的性能提升与有效应用。作为目前最大规模的3D人体数据集,我们期望其公开发布能推动人体中心任务的大规模创新。数据集已开放获取:https://kevinlee09.github.io/research/MVHumanNet++/
原文摘要 · Abstract (English)
In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while significant progress has been achieved in object-centric tasks through large-scale datasets like Objaverse and MVImgNet, human-centric tasks have seen limited advancement, largely due to the absence of a comparable large-scale human dataset. To bridge this gap, we present MVHumanNet++, a dataset that comprises multi-view human action sequences of 4,500 human identities. The primary focus of our work is on collecting human data that features a large number of diverse identities and everyday clothing using multi-view human capture systems, which facilitates easily scalable data collection. Our dataset contains 9,000 daily outfits, 60,000 motion sequences and 645 million frames with extensive annotations, including human masks, camera parameters, 2D and 3D keypoints, SMPL/SMPLX parameters, and corresponding textual descriptions. Additionally, the proposed MVHumanNet++ dataset is enhanced with newly processed normal maps and depth maps, significantly expanding its applicability and utility for advanced human-centric research. To explore the potential of our proposed MVHumanNet++ dataset in various 2D and 3D visual tasks, we conducted several pilot studies to demonstrate the performance improvements and effective applications enabled by the scale provided by MVHumanNet++. As the current largest-scale 3D human dataset, we hope that the release of MVHumanNet++ dataset with annotations will foster further innovations in the domain of 3D human-centric tasks at scale. MVHumanNet++ is publicly available at https://kevinlee09.github.io/research/MVHumanNet++/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。