用AI预测用户搜索词,提升视觉平台在生成式搜索中的流量
Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth
- 用视觉语言模型预测用户真实搜索意图,反向优化内容索引
- 构建可检索的图文聚合页,实现20%自然流量增长
- 适合想在生成式搜索时代提升曝光的视觉内容平台
大型语言模型正通过AI原生搜索系统重塑内容发现方式,如ChatGPT、Gemini和Claude。与传统关键词匹配不同,这些系统能推断用户意图、融合多模态证据并直接生成上下文答案,推动从搜索引擎优化(SEO)向生成式引擎优化(GEO)转变。对于拥有数十亿视觉资产的平台而言,单个图像缺乏生成式搜索所需的语义深度和权威信号,可能导致用户需求在页面内被满足而不再访问平台。本文提出Pinterest GEO,一个规模化框架,采用逆向搜索设计:不生成通用图像描述,而是微调视觉语言模型(VLMs)以预测用户实际会搜索的内容,并结合AI代理实时挖掘网络趋势,捕捉新兴搜索需求。这些由VLM生成的查询驱动构建基于多模态嵌入的语义连贯集合页,形成可索引的聚合内容,优化生成式检索。最后,采用混合VLM与双塔近似最近邻(ANN)架构,建立跨数十亿视觉资产的权威感知链接结构。该系统已部署于数十亿图像和数千万集合页,带来20%的自然流量增长,推动月活跃用户数实现数百万级增长,为视觉平台在生成式搜索时代的发展提供可落地的路径。
原文摘要 · Abstract (English)
Large Language Models are fundamentally reshaping content discovery through AI-native search systems such as ChatGPT, Gemini, and Claude. Unlike traditional search engines that match keywords to documents, these systems infer user intent, synthesize multimodal evidence, and generate contextual answers directly on the search page, introducing a paradigm shift from Search Engine Optimization (SEO) to Generative Engine Optimization (GEO). For visual content platforms hosting billions of assets, this poses an acute challenge: individual images lack the semantic depth and authority signals that generative search prioritizes, risking disintermediation as user needs are satisfied in-place without site visits. We present Pinterest GEO, a production-scale framework that pioneers reverse search design: rather than generating generic image captions describing what content is, we fine-tune Vision-Language Models (VLMs) to predict what users would actually search for, augmented this with AI agents that mine real-time internet trends to capture emerging search demand. These VLM-generated queries then drive construction of semantically coherent Collection Pages via multimodal embeddings, creating indexable aggregations optimized for generative retrieval. Finally, we employ hybrid VLM and two-tower ANN architectures to build authority-aware interlinking structures that propagate signals across billions of visual assets. Deployed at scale across billions of images and tens of millions of collections, GEO delivers 20\% organic traffic growth contributing to multi-million monthly active user (MAU) growth, demonstrating a principled pathway for visual platforms to thrive in the generative search era.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。