arXiv:2505.20292cs.CVcs.AI2025-05NeurIPS被引 68

构建百万级视频生成数据集与细粒度评测基准,提升主体一致性与自然度。

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

  • 设计细粒度评测框架,聚焦主体一致性和身份保真度。
  • 推出首个百万级高质量三元组数据集,支持多视角合成与跨视频关联。
  • 提出三种自动指标,精准衡量一致性、自然度和文本相关性,适合研究者使用。

主体到视频(S2V)生成旨在创建忠实反映参考内容的视频,提升视频制作灵活性。为建立S2V生成基础设施,本文提出OpenS2V-Nexus,包含(i)OpenS2V-Eval细粒度评测基准,及(ii)OpenS2V-5M百万规模数据集。与现有基于VBench的全局粗粒度评测不同,OpenS2V-Eval关注模型生成主体一致、外观自然、身份保真的能力。该基准引入7类共180个S2V提示,涵盖真实与合成测试数据,并提出NexusScore、NaturalScore和GmeScore三个自动指标,分别量化主体一致性、自然度和文本相关性。在此基础上,对18个代表性S2V模型进行综合评估,揭示其在不同内容下的优劣。同时,构建首个开源大规模S2V数据集OpenS2V-5M,包含五百万条高质量720P主体-文本-视频三元组。通过跨视频关联实现主体分割与配对,利用GPT-Image-1对原始帧生成多视角表示,确保主体信息多样性。OpenS2V-Nexus为未来S2V研究提供坚实基础设施。

原文摘要 · Abstract (English)

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus, consisting of (i) OpenS2V-Eval, a fine-grained benchmark, and (ii) OpenS2V-5M, a million-scale dataset. In contrast to existing S2V benchmarks inherited from VBench that focus on global and coarse-grained assessment of generated videos, OpenS2V-Eval focuses on the model's ability to generate subject-consistent videos with natural subject appearance and identity fidelity. For these purposes, OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data. Furthermore, to accurately align human preferences with S2V benchmarks, we propose three automatic metrics, NexusScore, NaturalScore and GmeScore, to separately quantify subject consistency, naturalness, and text relevance in generated videos. Building on this, we conduct a comprehensive evaluation of 18 representative S2V models, highlighting their strengths and weaknesses across different content. Moreover, we create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples. Specifically, we ensure subject-information diversity in our dataset by (1) segmenting subjects and building pairing information via cross-video associations and (2) prompting GPT-Image-1 on raw frames to synthesize multi-view representations. Through OpenS2V-Nexus, we deliver a robust infrastructure to accelerate future S2V generation research.

视频生成主体一致数据集评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。