构建首个面向游戏内容生成的多模态多视角数据集,支持超分辨率与可控视频生成研究。
$\mathtt{M^3VIR}$: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation
- 基于虚幻引擎5渲染80个场景,提供真实低清/高清配对与多视角图像
- 包含8大类游戏内容,涵盖超分辨率、新视角合成及联合任务,支持可控生成
- 首次提供对象级多风格真值数据集,适合云游戏与生成式AI研究者使用
游戏与娱乐产业正迅速发展,依赖沉浸式体验与生成式AI技术。训练此类模型需大规模、多样化的数据集来捕捉游戏环境的真实特性。然而现有数据集通常局限于特定领域或依赖人工退化,无法准确反映游戏内容的独特性。此外,可控视频生成尚无基准。为此,我们提出 $ exttt{M^3VIR}$,一个大规模、多模态、多视角数据集,旨在克服现有资源的局限。$ exttt{M^3VIR}$ 使用虚幻引擎5渲染80个场景,覆盖8个类别,提供高保真真实低清-高清配对图像与多视角帧。其子集 $ exttt{M^3VIR extunderscore MR}$ 支持超分辨率(SR)、新视角合成(NVS)及联合任务;$ exttt{M^3VIR extunderscore MS}$ 是首个对象级、多风格真值数据集,推动可控视频生成研究。我们还对多个先进SR与NVS方法进行基准测试,建立性能基线。尽管当前无方法直接处理可控视频生成,$ exttt{M^3VIR}$ 为该领域提供了重要基准。通过发布该数据集,我们希望促进下一代云游戏与娱乐中人工智能驱动的内容恢复、压缩与可控生成研究。
原文摘要 · Abstract (English)
The gaming and entertainment industry is rapidly evolving, driven by immersive experiences and the integration of generative AI (GAI) technologies. Training such models effectively requires large-scale datasets that capture the diversity and context of gaming environments. However, existing datasets are often limited to specific domains or rely on artificial degradations, which do not accurately capture the unique characteristics of gaming content. Moreover, benchmarks for controllable video generation remain absent. To address these limitations, we introduce $\mathtt{M^3VIR}$, a large-scale, multi-modal, multi-view dataset specifically designed to overcome the shortcomings of current resources. Unlike existing datasets, $\mathtt{M^3VIR}$ provides diverse, high-fidelity gaming content rendered with Unreal Engine 5, offering authentic ground-truth LR-HR paired and multi-view frames across 80 scenes in 8 categories. It includes $\mathtt{M^3VIR\_MR}$ for super-resolution (SR), novel view synthesis (NVS), and combined NVS+SR tasks, and $\mathtt{M^3VIR\_{MS}}$, the first multi-style, object-level ground-truth set enabling research on controlled video generation. Additionally, we benchmark several state-of-the-art SR and NVS methods to establish performance baselines. While no existing approaches directly handle controlled video generation, $\mathtt{M^3VIR}$ provides a benchmark for advancing this area. By releasing the dataset, we aim to facilitate research in AI-powered restoration, compression, and controllable content generation for next-generation cloud gaming and entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。