arXiv:2510.19186cs.CL2025-10被引 1

现有对话评估方法难发现工具调用中的隐藏错误。

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

  • 构建了覆盖多种错误场景的合成对话基准TRACE。
  • 主流评估框架在新基准上表现仍不理想,误差率高。
  • 适合研究对话系统鲁棒性与用户体验的学者参考。

评估使用外部工具的对话AI系统极具挑战性,因为用户、智能体和工具之间的复杂交互可能导致错误。现有评估方法仅关注用户满意度或智能体的工具调用能力,却无法捕捉多轮工具增强对话中的关键错误——例如智能体误读工具结果,但用户仍感到满意。为此,我们提出了TRACE,一个系统化合成的工具增强对话基准,涵盖多样化的错误案例。对当前最先进的对话评估框架进行评估发现,所有方法在该基准上表现均未达到理想水平,凸显了该任务的根本难度。

原文摘要 · Abstract (English)

Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and tools. While existing evaluation methods assess either user satisfaction or agents' tool-calling capabilities, they fail to capture critical errors in multi-turn tool-augmented dialogues-such as when agents misinterpret tool results yet appear satisfactory to users. We introduce TRACE, a benchmark of systematically synthesized tool-augmented conversations covering diverse error cases. Evaluation with state-of-the-art conversation evaluation frameworks reveals that all approaches remain far from ideal performance, demonstrating the fundamental difficulty of this benchmark.

对话评估工具调用基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。