arXiv:2601.14706cs.CV2026-01被引 5

打造实时更新的时尚图像检索基准,模拟真实电商场景。

LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval

  • 融合真实与AI生成时尚图,按时间戳动态更新。
  • 强模型在该基准上Recall@1普遍低于60%。
  • 适合关注电商视觉搜索与模型泛化能力的研究者。

本文提出LookBench,一个面向真实电商场景的实时、全面且具有挑战性的时尚图像检索基准。其包含来自实时网站的最新商品图与AI生成的时尚图像,反映当前潮流与使用场景。每个测试样本均带时间戳,支持定期更新,实现与训练截止时间对齐的污染感知评估。基于细粒度属性分类体系,覆盖单品与成套服饰检索。实验表明,该基准对主流模型构成显著挑战,多数模型Recall@1低于60%。自研模型表现最优,开源版本排名第二,两者在经典Fashion200K数据集上也达到领先水平。LookBench将每半年更新一次新样本及更难任务变体,提供可持续的性能衡量标准。我们公开了排行榜、数据集、评估代码与训练模型。

原文摘要 · Abstract (English)

In this paper, we present LookBench (We use the term "look" to reflect retrieval that mirrors how people shop -- finding the exact item, a close substitute, or a visually consistent alternative.), a live, holistic and challenging benchmark for fashion image retrieval in real e-commerce settings. LookBench includes both recent product images sourced from live websites and AI-generated fashion images, reflecting contemporary trends and use cases. Each test sample is time-stamped and we intend to update the benchmark periodically, enabling contamination-aware evaluation aligned with declared training cutoffs. Grounded in our fine-grained attribute taxonomy, LookBench covers single-item and outfit-level retrieval across. Our experiments reveal that LookBench poses a significant challenge on strong baselines, with many models achieving below $60\%$ Recall@1. Our proprietary model achieves the best performance on LookBench, and we release an open-source counterpart that ranks second, with both models attaining state-of-the-art results on legacy Fashion200K evaluations. LookBench is designed to be updated semi-annually with new test samples and progressively harder task variants, providing a durable measure of progress. We publicly release our leaderboard, dataset, evaluation code, and trained models.

图像检索时尚电商基准测试AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。