首个超高清图像编辑数据集,让高分辨率图片精准修改成为可能
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

- 构建120K组超清图像编辑三元组,每张图超过4096×4096像素
- 用多阶段过滤确保图像质量与指令一致性,提升细节还原度
- 适合做高分辨率图像生成、编辑的开发者和研究者使用
直接编辑超高清(UHR)图像具有重要价值但研究不足,主要受限于高质量数据缺失和高频纹理建模难题。我们提出VINS-120K,首个基于指令的超高清图像编辑大规模数据集,包含120,000组精心筛选的三元组:指令、输入图像与编辑后图像。每张图像分辨率均超过4K(≥4096×4096),并通过多阶段严格过滤流程,确保视觉质量、指令对齐性与美学保真度。基于该数据集,我们设计高频感知后适配策略,将预训练非高分辨率模型扩展至超高清场景。同时提出VINS-4KEval基准,涵盖多种编辑类型,实现超高清环境下的统一评估。实验表明,该方法显著提升细粒度细节生成与纹理真实感。
原文摘要 · Abstract (English)
Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency texture details. We introduce VINS-120K, the first large-scale dataset for instruction-based UHR image editing, comprising 120K carefully curated triplets of instruction, input image, and edited image. Each image exceeds 4K resolution ($\geq$4096 $\times$ 4096) and is filtered through a rigorous multi-stage pipeline to ensure visual quality, instruction alignment, and aesthetic fidelity. Built on VINS-120K, we further develop a high-frequency-aware post-adaptation strategy to extend pretrained non-high-resolution models to the UHR regime. We also present VINS-4KEval, a benchmark covering diverse editing types, to facilitate consistent evaluation in UHR settings. Experiments confirm that our work improves fine-grained detail synthesis and texture realism in UHR image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。