arXiv:2606.12924cs.AI2026-06被引 1

用模拟买家测试电商搜索智能体,发现记忆机制和模型选择影响显著。

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce

  • 设计双代理仿真框架,固定买家角色对比不同响应器性能。
  • 滚动窗口记忆比意图提取快35%且质量更优,修复后失败率降62%。
  • 不同大模型表现差异明显,评委评价标准也存在根本分歧。

我们提出一个模块化的双代理仿真框架,用于评估对话式购物助手架构。独立的买家代理配置了多种人物画像、任务目标和耐心水平,与可替换的响应器配对,后者接入真实电商搜索API。在实验中保持买家不变,实现对响应器设计的受控比较。基于14个人物画像桶中的2011次对话,得出四项实证发现:第一,滚动窗口记忆在所有质量指标上均优于意图提取记忆,且每查询速度提升35%;第二,通过系统性故障分析,针对某版本响应器实施针对性修复,使全数据集的失败和近失败率降低62%;第三,将响应器大模型从Gemini 2.5更换为Llama 3.3 70B,尽管架构一致,仍导致得分下降0.16–0.45分;第四,前沿大模型评审者间存在系统性哲学分歧:Gemini强调流程正确性,Claude则更关注具体结果,即使使用相同评估提示。

原文摘要 · Abstract (English)

We present a modular two-agent simulation framework for evaluating conversational shopping assistant architectures. An independent buyer agent, configured with personas, missions, and patience levels, is paired with an interchangeable responder that integrates with a real e-commerce search API. Holding the buyer constant across experiments enables controlled comparison of responder designs on identical scenarios. Using 2011 conversations across 14 persona buckets, we establish four empirical findings. First, rolling-window memory outperforms intent-extraction memory on all quality metrics while being 35% faster per query. Second, illustrating rapid evidence-driven iteration, a systematic failure analysis of a responder version enables targeted fixes that reduce failure and near-failure rates by 62% across the full dataset. Third, swapping the responder LLM backbone from Gemini~2.5 to Llama~3.3~70B costs 0.16--0.45 points despite identical architecture. Finally, we document systematic philosophical disagreement between frontier LLM judges: Gemini rewards process correctness while Claude demands concrete outcomes, despite using the same evaluation prompt.

智能体电商搜索大模型评估仿真框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。