arXiv:2605.20147cs.CV2026-05被引 2

构建10000万像素级图文数据集,推动文本生成超高清图像技术突破。

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

论文配图:PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
图 1 · 摘自论文原文
  • 设计高质量数据流水线,构建95000张百万像素以上图像数据集
  • 实现多类基础模型原生生成10000万像素图像,支持三种训练方案
  • 提出基于大语言模型的评估基准,系统评测超清图像质量与语义一致性

文本到图像(T2I)模型近年来在1K和2K分辨率上取得显著进展。随着对更佳视觉体验的需求及成像技术的快速发展,超高清(UHR)图像生成需求日益增长。然而,由于高分辨率内容稀缺且复杂,UHR图像生成面临巨大挑战。本文首次推出PixVerve-95K,一个高质量、开源的UHR T2I数据集,采用精心设计的数据处理流程,包含95,000张覆盖多样场景的图像(每张图像最小像素数达100M),并配有七维标注。基于该大规模图文数据集,我们首次实现多种T2I基础模型原生支持100MP图像生成,探索了三种训练策略。最后,结合传统指标与基于多模态大语言模型的评估方法,提出PixVerve-Bench基准,建立涵盖视觉质量与语义对齐的综合评价体系。大量实验结果与训练策略的深入分析为未来突破提供了宝贵洞见。

原文摘要 · Abstract (English)

Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imaging technology, the demand for Ultra-High-Resolution (UHR) image generation has grown significantly. However, UHR image generation poses great challenges due to the scarcity and complexity of high-resolution content. In this paper, we first introduce PixVerve-95K, a high-quality, open-source UHR T2I dataset curated with a carefully designed data pipeline, which contains 95K images across diverse scenarios (each image has a minimum pixel-count of 100M) and seven-dimensional annotations. Based on our large-scale image-text dataset, we take a pioneering step to extend various T2I foundation models to native 100MP generation with three training schemes. Finally, leveraging both conventional metrics and multimodal large language model-based assessments, our proposed PixVerve-Bench benchmark establishes a comprehensive evaluation protocol for UHR images encompassing visual quality and semantic alignment. Extensive experimental results on our benchmark and the constructive exploration of training strategies collaboratively provide valuable insights for future breakthroughs.

超高清生成图文数据集100MPT2I

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。