让服装图片支持无缝缩放,细节清晰且无需对齐。
GarmentZoom: Generating Zoomable Images from Garment Listings

- 用单模型实现跨服装的连续缩放,不需逐件微调。
- 在3-20倍缩放下保持细节质量,接近专用模型效果。
- 适合电商图像增强,提升用户浏览体验。
在线服装商品展示常包含一张全景图和一张局部特写图,但两者分别侧重视野范围与细节呈现,导致用户需频繁切换视图,打断浏览流程。我们提出GarmentZoom,通过增强全景图使其细节达到特写水平,实现无缝缩放与平移浏览。不同于传统基于参考图的超分辨率方法,本场景中的特写图与全景图空间未对齐,且缩放倍数差异大(3-20×)。现有方法多依赖对齐或需针对每件服装微调。我们训练一个统一模型,支持多样化服装的连续缩放范围,无需空间对齐即可合成细节,性能接近专用模型,训练成本仅为后者的极小部分。
原文摘要 · Abstract (English)
Online product listings for garments often include an overview photo and a close-up to show garment details. However, each photo focuses on either field of view or garment detail, forcing users to alternate between views and breaking browsing continuity. We present GarmentZoom, a system that enhances the full-view photo to match the fidelity of its accompanying close-up, enabling seamless zoom-and-pan exploration. Unlike standard reference-based super-resolution, our setting involves close-up references that are spatially unaligned with the full view, and scale factors that vary substantially across garments 3-20$\times$. Prior work typically relies on alignment to transfer details or requires per-instance fine-tuning to memorize them. Instead, we train a single model that supports a continuous range of scales across diverse garments. Our approach synthesizes details without requiring spatial alignment and matches the quality of per-instance methods with a fraction of the training cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。