打造实时更新的时尚图像检索基准,模拟真实电商场景。
LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval
- 融合真实与AI生成时尚图,按时间戳动态更新。
- 强模型在该基准上Recall@1普遍低于60%。
- 适合关注电商视觉搜索与模型泛化能力的研究者。
本文提出LookBench,一个面向真实电商场景的实时、全面且具有挑战性的时尚图像检索基准。其包含来自实时网站的最新商品图与AI生成的时尚图像,反映当前潮流与使用场景。每个测试样本均带时间戳,支持定期更新,实现与训练截止时间对齐的污染感知评估。基于细粒度属性分类体系,覆盖单品与成套服饰检索。实验表明,该基准对主流模型构成显著挑战,多数模型Recall@1低于60%。自研模型表现最优,开源版本排名第二,两者在经典Fashion200K数据集上也达到领先水平。LookBench将每半年更新一次新样本及更难任务变体,提供可持续的性能衡量标准。我们公开了排行榜、数据集、评估代码与训练模型。
原文摘要 · Abstract (English)
In this paper, we present LookBench (We use the term "look" to reflect retrieval that mirrors how people shop -- finding the exact item, a close substitute, or a visually consistent alternative.), a live, holistic and challenging benchmark for fashion image retrieval in real e-commerce settings. LookBench includes both recent product images sourced from live websites and AI-generated fashion images, reflecting contemporary trends and use cases. Each test sample is time-stamped and we intend to update the benchmark periodically, enabling contamination-aware evaluation aligned with declared training cutoffs. Grounded in our fine-grained attribute taxonomy, LookBench covers single-item and outfit-level retrieval across. Our experiments reveal that LookBench poses a significant challenge on strong baselines, with many models achieving below $60\%$ Recall@1. Our proprietary model achieves the best performance on LookBench, and we release an open-source counterpart that ranks second, with both models attaining state-of-the-art results on legacy Fashion200K evaluations. LookBench is designed to be updated semi-annually with new test samples and progressively harder task variants, providing a durable measure of progress. We publicly release our leaderboard, dataset, evaluation code, and trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。