arXiv:2606.27708cs.CV2026-06

用知识蒸馏和权重插值,让视觉语言模型在时装检索上既精准又泛化。

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

  • 全量微调+领域数据蒸馏,再与基座模型加权融合。
  • 在多个基准上超越LoRA、更大模型和外部数据训练方案。
  • 新发布高质量时尚检索数据集,揭示并修正现有数据偏见。

将基础视觉语言编码器适配至特定检索任务时存在核心矛盾:目标分布上的性能提升会损害模型的广泛泛化能力,而时尚检索正是这一问题的严苛案例。我们提出ZooClaw-FashionSigLIP2——一个专用于时尚检索的SigLIP2-base模型,通过简单有效的方案解决该矛盾:在精选领域数据上进行全量微调并结合知识蒸馏,随后采用wiseft权重插值融合基座模型。该方法在公平评估下,在所有评测基准中均优于现有基线,包括LoRA、参数高达10亿的更大骨干网络以及外部训练数据。此外,我们发布了ZooClaw-Fashion——一个高质量的时尚检索新基准,并对常用基准进行了系统性质量分析,识别并缓解了其公开真值中的结构性偏差。模型权重及全部评估工具均已开源,以促进后续研究。

原文摘要 · Abstract (English)

Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains on the target distribution come at the cost of the foundation model's broad generalization, and fashion retrieval is a stringent instance of this problem. We present ZooClaw-FashionSigLIP2, a fashion-specialized SigLIP2-base model that resolves this tradeoff with a simple recipe -- full fine-tuning with knowledge distillation on curated in-domain data, followed by \wiseft~\citep{wortsman2022wiseft} weight interpolation with the base model -- and outperforms LoRA, larger backbones (up to 1B parameters), and external training data. Under fair evaluation, ZooClaw-FashionSigLIP2 outperforms all baselines on every benchmark in our suite. In addition, we release ZooClaw-Fashion, a new high-quality fashion retrieval benchmark, and a systematic quality analysis of widely-used benchmarks that exposes and mitigates structural biases in their public ground truth. We open-source the model weights and all evaluation artifacts to facilitate future research.

时尚检索知识蒸馏视觉语言模型模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。