arXiv:2502.18308cs.CL2025-02被引 4

用AI代理动态测试大模型如何回应反驳,发现其难记信息且越辩越差。

RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction

  • 引入AI代理作为反驳者和评估者,实现灵活多样的反馈测试
  • 模型能响应反驳但无法长期记住反驳内容,且任务性能随反驳增多下降
  • 揭示大模型在长对话中保留与使用历史信息的缺陷,适合研究对话鲁棒性者关注

在多轮交互场景中,大语言模型可利用用户反馈提升回复质量与相关性。然而,评估模型对用户反驳反馈的处理能力仍具挑战。本文提出RefuteBench 2.0,显著扩展原版基准,引入大模型代理作为反驳者与评估者,实现灵活全面的评估。设计了瞬时与持久两类具有不同有效期限的反驳指令。元评估显示,基于LLM的反驳者生成更类人化的反驳,评估者评分与人类高度相关。多种大模型实验表明,当前模型虽能有效响应反驳,但难以记忆反驳信息;有趣的是,随着反驳次数增加,初始任务性能反而下降。注意力分析进一步揭示当前模型在长上下文对话中保留并正确使用先前信息存在潜在弱点。

原文摘要 · Abstract (English)

In the multi-turn interaction schema, large language models (LLMs) can leverage user feedback to enhance the quality and relevance of their responses. However, evaluating an LLM's ability to incorporate user refutation feedback is crucial yet challenging. In this study, we introduce RefuteBench 2.0, which significantly extends the original RefuteBench by incorporating LLM agents as refuters and evaluators, which allows for flexible and comprehensive assessment. We design both transient and persistent refutation instructions with different validity periods. Meta-evaluation shows that the LLM-based refuter could generate more human-like refutations and the evaluators could assign scores with high correlation with humans. Experimental results of various LLMs show that current models could effectively satisfy the refutation but fail to memorize the refutation information. Interestingly, we also observe that the performance of the initial task decreases as the refutations increase. Analysis of the attention scores further shows a potential weakness of current LLMs: they struggle to retain and correctly use previous information during long context dialogues. https://github.com/ElliottYan/RefuteBench-2.0

大模型评测对话系统反驳测试注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。