arXiv:2511.21389cs.IRcs.AI2025-11

用注意力引导的视觉文本结构,精准识别重复商品

FITRep: Attention-Guided Item Representation via MLLMs

  • 基于视觉文本结构分层提取概念,避免信息丢失
  • 在线测试提升3.6%点击率和4.25%广告单价
  • 适合需要去重的电商与广告平台使用

在线平台常因视觉与文字相似的近似重复商品导致用户体验下降。尽管多模态大模型可生成多模态嵌入,但现有方法将表示视为黑箱,忽略结构关系(如主次元素),引发局部结构坍塌问题。受特征整合理论启发,我们提出FITRep,首个面向细粒度商品去重的注意力引导白盒表示框架。包含:(1) 概念分层信息提取(CHIE),利用MLLMs提取层次化语义概念;(2) 结构保持降维(SPDR),基于自适应UMAP实现高效信息压缩;(3) FAISS聚类(FBC),通过FAISS为每项分配唯一聚类ID。在美团广告系统部署中,线上A/B测试显示CTR提升3.60%,CPM提升4.25%,验证了其有效性和实际应用价值。

原文摘要 · Abstract (English)

Online platforms usually suffer from user experience degradation due to near-duplicate items with similar visuals and text. While Multimodal Large Language Models (MLLMs) enable multimodal embedding, existing methods treat representations as black boxes, ignoring structural relationships (e.g., primary vs. auxiliary elements), leading to local structural collapse problem. To address this, inspired by Feature Integration Theory (FIT), we propose FITRep, the first attention-guided, white-box item representation framework for fine-grained item deduplication. FITRep consists of: (1) Concept Hierarchical Information Extraction (CHIE), using MLLMs to extract hierarchical semantic concepts; (2) Structure-Preserving Dimensionality Reduction (SPDR), an adaptive UMAP-based method for efficient information compression; and (3) FAISS-Based Clustering (FBC), a FAISS-based clustering that assigns each item a unique cluster id using FAISS. Deployed on Meituan's advertising system, FITRep achieves +3.60% CTR and +4.25% CPM gains in online A/B tests, demonstrating both effectiveness and real-world impact.

去重多模态广告系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。