arXiv:2605.09918cs.LGcs.AI2026-05被引 2

首个面向大模型广告的综合性数据集,助力平衡用户体验与商业收益。

NaiAD: Initiate Data-Driven Research for LLM Advertising

论文配图:NaiAD: Initiate Data-Driven Research for LLM Advertising
图 1 · 摘自论文原文
  • 构建5.9万条带广告的对话数据,分离评估用户与商业价值。
  • 提出解耦生成流程,实现广告融入策略多样化。
  • 揭示四种有效广告融合语义策略,支持灵活控制目标。

在大模型广告中协调平台收益与用户体验,需以数据为核心。本文提出NaiAD,首个面向大模型原生广告的综合性数据集,包含58,999条精心构造的、嵌入广告的回复与用户查询对。NaiAD基于理论支撑的评估指标,分别且全面捕捉用户与商业效用。为缓解对齐大模型的维度共线性问题,我们设计了一种解耦生成管道,生成结构多样化的样本,涵盖显式分离利益相关方效用、各维度整体强或弱的响应。我们进一步通过方差校准预测驱动推理(VC-PPI)框架提供校准后的评分标签,使自动评分与人工标注一致。机制分析显示,成功的广告融合依赖于四种不同的语义策略。利用NaiAD训练的模型可内化这些策略,同时提升用户与商业效用,并通过上下文学习独立控制不同目标。这些成果使NaiAD成为未来大模型原生广告系统开发的基础基础设施。

原文摘要 · Abstract (English)

Reconciling platform revenue with user experience in LLM advertising motivates a data-centric foundation. We introduce NaiAD, the first comprehensive dataset for LLM-native advertising comprising 58,999 carefully constructed ad-embedded responses paired with user queries. NaiAD is organized around theoretically grounded evaluation metrics that separately and comprehensively capture user and commercial utility. To mitigate the dimensional collinearity of aligned LLMs, we propose a decoupled generation pipeline that produces structurally diverse samples, ranging from responses that explicitly disentangle stakeholder utilities to responses that are uniformly strong or weak across dimensions. We further provide score labels calibrated by a Variance-Calibrated Prediction-Powered Inference (VC-PPI) framework, aligning automated scoring with human annotations. Mechanistic analyses reveal that successful ad integration relies on reasoning paths that cluster into four distinct semantic strategies. Models leveraging NaiAD internalize these strategies to simultaneously improve user and commercial utility, while enabling independent control over these distinct objectives via in-context learning. Together, these results position NaiAD as a foundational infrastructure for developing future LLM-native ad systems.

大模型广告数据集用户-商业平衡生成策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。