arXiv:2505.23662cs.CL2025-05EMNLP被引 5

测试大模型在长期对话中使用工具的稳定性

ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions

  • 设计连续对话场景,模拟真实使用中的任务流与干扰
  • 14个主流大模型在长程任务中表现明显下降
  • 适合评估模型在复杂、持续交互下的实际应用能力

大型语言模型(LLMs)在使用外部工具解决用户问题方面展现出强大能力。然而,现有评估大多假设工具使用发生在短上下文环境中,难以揭示模型在真实长期交互中的行为。为填补这一空白,我们提出ToolHaystack,一个用于测试长时交互中工具使用能力的基准。每个测试实例包含多个任务执行上下文和连续对话中的真实噪声,可评估模型维持上下文与应对各种干扰的能力。将该基准应用于14个最先进的LLM,发现尽管模型在标准多轮设置中表现良好,但在ToolHaystack中常出现显著退化,暴露出此前工具评测未揭示的关键长期鲁棒性缺陷。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capabilities in long-term interactions. Each test instance in ToolHaystack includes multiple tasks execution contexts and realistic noise within a continuous conversation, enabling assessment of how well models maintain context and handle various disruptions. By applying this benchmark to 14 state-of-the-art LLMs, we find that while current models perform well in standard multi-turn settings, they often significantly struggle in ToolHaystack, highlighting critical gaps in their long-term robustness not revealed by previous tool benchmarks.

大模型评估工具调用长程交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。