arXiv:2504.07567cs.CVcs.AI2025-04中稿 · Future Technologie…被引 2

对比多种模型在电商图像任务中的表现,给出高效实用的调优建议。

Benchmarking Image Embeddings for E-Commerce: Evaluating Off-the Shelf Foundation Models, Fine-Tuning Strategies and Practical Trade-offs

  • 测试了卷积与变压器模型在监督、自监督和图文对比学习下的嵌入效果
  • 全量微调性能最好,但顶部微调可显著降低计算开销且效果接近
  • 根据数据集特性选择策略,适合电商场景的模型选型与优化

我们对电商场景中的图像嵌入进行了基准测试,评估预训练卷积与变压器模型在分类和检索任务中的适用性。涵盖通过监督、自监督及图文对比学习训练的嵌入方法。在六类电商数据集(时尚、消费品、汽车、食品、零售)上,对比全量微调与仅顶部微调的迁移学习策略。结果表明,全量微调表现稳定,而图文对比与自监督嵌入在较少训练下即可达到相近效果;监督嵌入跨架构表现一致,自监督与对比嵌入则差异较大,通常受益于顶部微调。顶部微调成为全量微调的有效替代,大幅降低计算成本。还探索了跨域微调,其效果取决于数据集特征。研究为电商应用提供了兼顾效率与性能的嵌入选择与调优指导。

原文摘要 · Abstract (English)

We benchmark foundation models image embeddings for classification and retrieval in e-Commerce, evaluating their suitability for real-world applications. Our study spans embeddings from pre-trained convolutional and transformer models trained via supervised, self-supervised, and text-image contrastive learning. We assess full fine-tuning and transfer learning (top-tuning) on six diverse e-Commerce datasets: fashion, consumer goods, cars, food, and retail. Results show full fine-tuning consistently performs well, while text-image and self-supervised embeddings can match its performance with less training. While supervised embeddings remain stable across architectures, SSL and contrastive embeddings vary significantly, often benefiting from top-tuning. Top-tuning emerges as an efficient alternative to full fine-tuning, reducing computational costs. We also explore cross-tuning, noting its impact depends on dataset characteristics. Our findings offer practical guidelines for embedding selection and fine-tuning strategies, balancing efficiency and performance.

图像嵌入电商微调策略模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。