arXiv:2608.30685cs.AI2026-08

提出双时域评估框架,精准定位工业级工具调用智能体的执行问题。

ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents

论文配图:ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents
图 1 · 摘自论文原文
  • 分请求与交互双时域诊断,定位错误发生位置与上下文偏离
  • 在美团小团团真实流量中提升用户参与度与业务结果
  • 可落地的评估信号支持策略优化,适合工业级AI服务迭代

大型语言模型(LLM)代理在需动态业务条件下迭代使用工具的用户服务中日益普及。可靠评估对持续改进至关重要:须揭示能力缺陷、指导优化优先级并评估干预效果。然而,工业级代理服务既体现在单个请求的迭代轨迹上,也贯穿于持续的用户交互过程。仅依赖最终结果评估可能掩盖缺陷根源,且无法判断后续服务是否仍与早期上下文一致。为此,我们提出ATLAS——一种面向工业级工具调用代理的双时域诊断评估框架。在请求时域,基于轨迹的诊断信号将缺陷关联至执行位置与能力问题;在交互时域,基于用户的信号评估服务在持续交互中是否保持响应性。二者结合提供结构化诊断证据,用于分析执行缺陷与持续服务行为。ATLAS将诊断信号实例化为可执行的接口,具备明确证据范围和决策边界。采用真实业务日志中的高置信参考校准LLM裁判接口;必要时,将其决策行为蒸馏为高效诊断模型,实现低延迟、低成本评估。由此产生的反馈支持策略优化。我们在美团小团团生产流量上评估ATLAS。离线实验验证诊断信号保真度与基于回放的策略改进效果,线上A/B实验显示用户参与度、下游业务指标与抽样人工审计质量同步提升。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essential for sustained improvement: it must reveal capability deficiencies, inform priorities, and assess interventions. Yet industrial agent service unfolds both through the iterative trajectory of a current request and through continued user interaction. Final-outcome assessment can therefore obscure where deficiencies arise and whether later service remains aligned with context from earlier exchanges. We propose ATLAS, a dual-horizon diagnostic evaluation framework for industrial tool-use agents. At the request horizon, trajectory-wise diagnostic signals relate deficiencies to execution locations and capability concerns. At the interaction horizon, user-wise signals assess whether service remains responsive across continued interaction. Together, these views provide structured diagnostic evidence for analyzing execution deficiencies and sustained service behavior. ATLAS instantiates them as executable signals with explicit evidence scopes and decision boundaries. LLM judge interfaces are calibrated against high-confidence references from real business logs; when needed, their decision behavior is distilled into efficient diagnostic models for lower-latency, lower-cost evaluation. The resulting feedback supports policy optimization. We evaluate ATLAS on Meituan Xiaotuan production traffic. Offline experiments assess diagnostic-signal fidelity and replay-based policy improvement, while online A/B experiments show concurrent gains in user engagement, downstream business outcomes, and sampled human-audit quality.

评估框架工具调用工业应用LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。