构建大规模人体视频生成数据集与评测体系,解决真实场景下人像生成难题
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

- 构建分层标注的多场景人体视频数据集,支持细粒度建模
- 提出三层次评测体系,指标与人类感知高度一致
- 适合研究视频生成、人机交互与多模态合成的开发者
近年来,音视频联合生成模型在内容创作方面展现出强大能力。然而,在复杂真实物理场景中生成高保真人体中心视频仍面临重大挑战。我们发现根源在于现有数据集在三个维度存在结构性缺陷:全局场景与摄像机视角多样性不足、人与人/物之间的交互建模稀疏、个体属性对齐不充分。为此,我们提出OmniHuman,一个大规模、多场景人体建模数据集。该数据集提供层级化标注,涵盖视频级场景、帧级交互和个体级属性。为支持数据采集与多模态标注,我们开发了全自动高质量数据收集流程。同时,我们建立OmniHuman基准(OHBench),一套三级评估体系,可科学诊断人体中心音视频合成性能。关键的是,OHBench引入与人类感知高度一致的指标,填补现有评测在全局场景、关系交互和个体属性层面的空白。
原文摘要 · Abstract (English)
Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a significant challenge. We identify that the root cause lies in the structural deficiencies of existing datasets across three dimensions: limited global scene and camera diversity, sparse interaction modeling (both person-person and person-object), and insufficient individual attribute alignment. To bridge these gaps, we present OmniHuman, a large-scale, multi-scene dataset designed for fine-grained human modeling. OmniHuman provides a hierarchical annotation covering video-level scenes, frame-level interactions, and individual-level attributes. To facilitate this, we develop a fully automated pipeline for high-quality data collection and multi-modal annotation. Complementary to the dataset, we establish the OmniHuman Benchmark (OHBench), a three-level evaluation system that provides a scientific diagnosis for human-centric audio-video synthesis. Crucially, OHBench introduces metrics that are highly consistent with human perception, filling the gaps in existing benchmarks by providing a comprehensive diagnosis across global scenes, relational interactions, and individual attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。