arXiv:2608.30023cs.IRcs.CL2026-08

构建百万级买家画像数据集,用于分析生成式引擎的推荐逻辑。

Demand-Side Measurement for Generative Engine Optimization: Constructing and Validating a Million-Persona, Intent-Annotated Buyer Corpus

  • 用合成数据+去重技术生成103万买家画像,覆盖511个行业和4个市场场景。
  • 每条画像含搜索查询、意图标签(信息/商业/交易)及信任来源偏好。
  • 专为研究生成式引擎如何选品设计,适合广告与推荐系统研究者使用。

生成式引擎如ChatGPT、Gemini和Perplexity可直接回答买家问题并列出品牌短名单。要研究品牌如何进入或未能进入该短名单,需需求端数据:买家在品类中问什么、需要何种信息、信任哪些来源。现有大规模人物画像数据集主要用于训练数据多样性,缺乏分阶段的搜索意图标签和首选来源字段,无法与供给端推荐结果关联。本文构建并验证了PersonaGen-1M,包含1,031,732条合成买家画像,覆盖511个行业标签和4个市场情境,共19,416,821个结构化行为属性,其中5,160,046条为搜索查询。每条画像带有单一主意图标签(78.3%信息性,17.4%商业性,4.3%交易性),以及一个标注买家信任来源类型的首选来源字段。数据源自约4000万条原始人物描述,经GPU加速的MinHash LSH与语义去重处理,后统一到固定模式。意图字段筛选出驱动推荐的商业评估型画像,首选来源字段可与引用溯源数据配对——此关联是主要用途,其控制性实证估计为未来工作。截至2026年8月,百万级画像数据集中仅有一份具备来源偏好属性,且仅为六值媒体渠道枚举;而PersonaGen-1M则为每条画像提供具体来源列表、分阶段意图标签与完整查询集合。完整数据集可申请非商业研究使用,分层子集已公开发布,供协议、模式与验证方法复用。

原文摘要 · Abstract (English)

Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying how brands enter or fail to enter that shortlist requires demand-side data: what buyers in a category ask, what information they need, and which sources they trust. Existing large persona corpora are built for training-data diversity and carry neither a staged search-intent label nor a preferred-sources field, so they cannot be joined to supply-side recommendation measurements. We built and validated PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, carrying 19,416,821 structured behavioral attributes, 5,160,046 of them search queries. Each persona carries a single primary_intent label covering its query set (78.3% informational, 17.4% commercial, 4.3% transactional) and a preferred_sources field naming the source types that buyer would trust. The corpus was built from roughly 40 million raw persona descriptions drawn from four public datasets through GPU-accelerated MinHash LSH plus semantic deduplication, then enriched to a fixed schema. The intent field selects the commercial-evaluation personas whose queries drive recommendation, and the preferred_sources field pairs against citation-provenance data; that join is the primary intended use, and its controlled empirical estimate is future work. Among million-scale persona corpora surveyed in August 2026, one other carries a source-preference attribute, as a six-value media-channel enum; PersonaGen-1M pairs named per-persona source lists with a staged commercial search-intent label and an attached query set. The full corpus is shared on request for non-commercial research; a stratified subset is published openly so the protocol, the schema and the validation can be inspected and reused without asking us.

生成式引擎买家画像推荐系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。