将文字攻击转为优势,把商品信息直接画在图上提升搜索准确率
Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval
- 把标题描述等文本直接渲染到商品图上,强化图文关联
- 在三类电商数据集上,图文检索准确率全面提升
- 适合做零样本电商多模态搜索的轻量级优化方案
电商平台的多模态商品检索系统依赖视觉与文本信号的有效融合以提升搜索相关性与用户体验。然而,如CLIP等视觉语言模型易受字体攻击影响,即图像中嵌入的误导性或无关文字会扭曲模型预测。本文提出一种新方法,逆向利用字体攻击逻辑:将商品标题、描述等关键文本直接渲染至商品图像上,实现视觉-文本压缩,从而增强图像与文本的对齐效果,提升多模态商品检索性能。我们在三个垂直领域电商数据集(运动鞋、手袋、收藏卡)上,使用六种主流视觉基础模型进行评估。实验表明,该方法在不同类别和模型族中均显著提升单模态与多模态检索准确率。结果表明,将商品元数据可视化地叠加于图像上,是一种简单但高效的零样本多模态检索增强策略。
原文摘要 · Abstract (English)
Multimodal product retrieval systems in e-commerce platforms rely on effectively combining visual and textual signals to improve search relevance and user experience. However, vision-language models such as CLIP are vulnerable to typographic attacks, where misleading or irrelevant text embedded in images skews model predictions. In this work, we propose a novel method that reverses the logic of typographic attacks by rendering relevant textual content (e.g., titles, descriptions) directly onto product images to perform vision-text compression, thereby strengthening image-text alignment and boosting multimodal product retrieval performance. We evaluate our method on three vertical-specific e-commerce datasets (sneakers, handbags, and trading cards) using six state-of-the-art vision foundation models. Our experiments demonstrate consistent improvements in unimodal and multimodal retrieval accuracy across categories and model families. Our findings suggest that visually rendering product metadata is a simple yet effective enhancement for zero-shot multimodal retrieval in e-commerce applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。