为大模型广告分析能力设计真实场景评测基准,解决答案过时难题。
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
- 基于真实广告平台请求构建评测集,动态重演专家操作生成最新答案。
- 最佳模型在复杂任务上准确率仅61.4%,暴露现有模型短板。
- 适合评估大模型在营销分析等专业领域的多轮协作能力。
尽管大语言模型代理在复杂推理方面取得显著进展,但在真实环境中的评估仍是一个开放问题。现有基准大多局限于理想化模拟,未能涵盖广告与营销分析等专业领域,这些领域需要与专业工具进行多轮交互,且答案随数据和平台规则变化迅速失效。为此,我们提出 AD-Bench,一个基于生产级广告平台真实用户营销分析请求构建的基准。AD-Bench 采用两项关键设计:(i) 动态真值流水线,通过重放专家工具调用轨迹生成与当前环境一致的答案,缓解答案过时问题;(ii) 轨迹感知评估,联合衡量端到端答案正确性(Pass@k)与轨迹覆盖度。请求按难度分为三个层级(L1-L3),用于检验多轮、多工具协作能力。实验表明,表现最优的模型 Claude-Opus-4.7 在整体上达到 Pass@1 = 76.9%、Pass@3 = 80.4%,轨迹覆盖率达 82.7%,但在最复杂任务 L3 上下降至 Pass@1 = 61.4%、Pass@3 = 65.1%,揭示即使顶尖代理在复杂广告分析中仍有明显不足。
原文摘要 · Abstract (English)
While Large Language Model (LLM) agents have made remarkable progress on complex reasoning, evaluating them in real-world environments remains an open problem. Existing benchmarks are largely confined to idealized simulations and fail to capture specialized domains such as advertising and marketing analytics, where tasks require multi-round interaction with professional tools and where ground-truth answers quickly become obsolete as data and platform rules evolve. To address this, we propose AD-Bench, a benchmark built from real user marketing-analysis requests on a production advertising platform. AD-Bench introduces two key designs: (i) a dynamic ground-truth pipeline that replays expert tool-call trajectories to regenerate answers consistent with the current environment, mitigating answer obsolescence; and (ii) a trajectory-aware evaluation that jointly measures end-to-end answer correctness (Pass@k) and trajectory coverage. Requests are stratified into three difficulty levels (L1-L3) to probe multi-round, multi-tool collaboration. Experiments show that the best model, Claude-Opus-4.7, attains Pass@1 = 76.9% and Pass@3 = 80.4% with 82.7% trajectory coverage overall, yet drops sharply on L3 to Pass@1 = 61.4% and Pass@3 = 65.1%, revealing that even state-of-the-art agents have substantial gaps in complex advertising analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。