arXiv:2608.26180cs.CLcs.AI2026-08

为个性化购物助手设计全新评估框架,解决推荐一致性与证据可追溯难题。

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

论文配图:PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
图 1 · 摘自论文原文
  • 提出 PACE 评估标准,涵盖个性化、可操作性、组合性与证据根基。
  • 构建含 2.26 万条结构化数据的 PACEShop 数据集,支持缺陷精确定位。
  • 开发无需训练的 PACEJudge 协议,提升多组件一致性与证据匹配度。

购物助手正从单纯排序转向结构化决策支持,需融合用户上下文、产品证据与下一步指导。现有评估仅覆盖部分问题,缺乏对结构化响应的联合评价目标。本文提出 PACE(个性化、可操作性、组合性、证据根基)评估体系,包含两个核心:PACEShop 基准数据集(22,625 条控制记录,含结构化人物画像、可审计证据池、良/差标签及黄金缺陷类别与位置标注)和 PACEJudge 无训练评判协议(通过结构化输出合约实现可报告性)。实验表明,通用判别器虽能识别整体质量,却无法还原 PACE 所需诊断字段;PACEShop 可验证这些失效,而 PACEJudge 在不重新训练前提下显著提升人物-来源对齐、跨组件一致性、证据锚定及缺陷类别/位置闭合能力,证明真实购物助手评估需任务匹配的输出合约,而非更强模型或标量提示。

原文摘要 · Abstract (English)

Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper context, contradicting itself across components, or leaving defects too vague to localize. Existing personalization, grounding, and LLM-as-a-judge benchmarks cover pieces of this problem, but they do not define a joint evaluation target for structured shopping-assistant responses. We formulate this missing evaluation target as PACE: Personalized, Actionable, Compositional, and Evidence-grounded evaluation. We instantiate PACE with two artifacts: PACEShop, a benchmark dataset that makes the target measurable through 22,625 controlled records with structured personas, auditable evidence pools, GOOD/BAD labels, and gold defect family and location annotations; and PACEJudge, a training-free judging protocol that makes the target reportable through a structured output contract. Our experiments show that generic judges can recognize broad quality but fail to recover the diagnostic fields required for PACE; PACEShop makes these failures verifiable, and PACEJudge improves persona-source, cross-component, grounding, and family/location closure without retraining, showing that realistic shopping-assistant evaluation requires a task-matched output contract rather than only a stronger backbone or scalar prompt.

评估基准购物助手LLM评测结构化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。