arXiv:2409.20222cs.CLcs.AI2024-09NeurIPS被引 22

提出动态对话评测框架,检验大模型在多任务并发下的长期记忆能力。

Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models

  • 设计长时、多任务交错的模拟对话流程
  • 短上下文模型+外部记忆系统表现优于大上下文模型
  • 揭示现有评测未捕捉的真实对话挑战

我们提出一种动态对话评测系统,通过一次模拟的、长时间的用户-代理交互来评估对话智能体的表现。该交互包含多个任务并行展开,通过频繁上下文切换实现任务交织,构建出真实场景,用以评估智能体的长期记忆(Long-Term Memory)、持续学习(Continual Learning)和信息整合(Information Integration)能力。对专有及开源大语言模型的测试结果表明,尽管大模型在单一任务中表现良好,但在任务交错时性能显著下降。值得注意的是,采用外部长期记忆(LTM)系统的短上下文模型表现与大上下文模型相当甚至更优。该评测揭示了当前基准无法捕捉的大模型应对自然对话的其他关键挑战。

原文摘要 · Abstract (English)

We introduce a dynamic benchmarking system for conversational agents that evaluates their performance through a single, simulated, and lengthy user$\leftrightarrow$agent interaction. The interaction is a conversation between the user and agent, where multiple tasks are introduced and then undertaken concurrently. We context switch regularly to interleave the tasks, which constructs a realistic testing scenario in which we assess the Long-Term Memory, Continual Learning, and Information Integration capabilities of the agents. Results from both proprietary and open-source Large-Language Models show that LLMs in general perform well on single-task interactions, but they struggle on the same tasks when they are interleaved. Notably, short-context LLMs supplemented with an LTM system perform as well as or better than those with larger contexts. Our benchmark suggests that there are other challenges for LLMs responding to more natural interactions that contemporary benchmarks have heretofore not been able to capture.

对话评测长期记忆多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。