arXiv:2502.05756cs.CVcs.LG2025-02被引 1

用ViT分析二手车零件图像,发现可识别可疑交易模式。

Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces

  • 用ViT提取图像嵌入,再通过UMAP降维和K均值聚类分析
  • 能有效分离相似零件图像,但存在重叠簇和异常点
  • 适合反欺诈研究者探索视觉模型在电商平台的应用

本研究探讨了视觉变换器(ViT)模型在生成来自Craigslist、OfferUp等在线市场平台的汽车零件图像视觉嵌入方面的性能。仅使用单模态数据,分析旨在检测可能涉及非法活动的视觉模式。流程包括从图像中提取高维嵌入,采用均匀流形近似与投影(UMAP)进行降维以可视化嵌入空间,并利用K-Means聚类对相似物品进行分类。每个聚类中心最近的代表性帖子揭示了聚类的组成与特征。结果表明,ViT在分离视觉模式方面具有优势,但聚类重叠和异常点仍暴露了单一模态方法在此领域的局限性。该工作有助于理解视觉变换器在在线市场分析中的作用,并为未来检测欺诈或非法行为提供基础。

原文摘要 · Abstract (English)

This study examines the capabilities of the Vision Transformer (ViT) model in generating visual embeddings for images of auto parts sourced from online marketplaces, such as Craigslist and OfferUp. By focusing exclusively on single-modality data, the analysis evaluates ViT's potential for detecting patterns indicative of illicit activities. The workflow involves extracting high-dimensional embeddings from images, applying dimensionality reduction techniques like Uniform Manifold Approximation and Projection (UMAP) to visualize the embedding space, and using K-Means clustering to categorize similar items. Representative posts nearest to each cluster centroid provide insights into the composition and characteristics of the clusters. While the results highlight the strengths of ViT in isolating visual patterns, challenges such as overlapping clusters and outliers underscore the limitations of single-modal approaches in this domain. This work contributes to understanding the role of Vision Transformers in analyzing online marketplaces and offers a foundation for future advancements in detecting fraudulent or illegal activities.

视觉嵌入ViT反欺诈图像聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。