新基准动态模拟用户对话,暴露大模型长对话中的真实短板
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
- 用三层追踪与生成代理模拟连续用户行为,更贴近真实交互
- 多轮对话中模型表现随深度下降,GPT-5保持66.40%鲁棒性最优
- 适合评估模型在复杂、长时对话中的稳定性与恢复能力
评估大模型在多主题对话中遵循指令的能力至关重要但极具挑战。现有基准仅限固定轮次,易饱和且忽视用户交互体验。本文提出一种新框架,包含三层追踪机制和查询生成代理,模拟连续用户行为。基于流理论,引入过程中心指标,在耗尽用户耐心时终止评估。基于此框架,我们构建了涵盖12类约束的动态演化基准EvolIF。分析发现模型在失败恢复和细粒度指令遵循上存在缺陷,性能分化随对话深度加剧。GPT-5表现出最强韧性,鲁棒性得分为66.40%,优于Gemini-3-Pro的60.81%,其他模型则显著落后。数据与代码将开源至https://github.com/JiaQiSJTU/EvolIF。
原文摘要 · Abstract (English)
Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to saturation and failing to account for users' interactive experience. In this work, we propose a novel framework featuring a three-layer tracking mechanism and a query synthesis agent to mimic sequential user behaviors. Grounded in Flow Theory, we introduce process-centric metrics and terminate a conversational evaluation only upon exhausting user patience. Leveraging this framework, we present EvolIF, an evolving benchmark covering 12 constraint groups. Our analysis reveals deficiencies in failure recovery and fine-grained instruction following, with performance stratification becoming evident as conversational depth increases. GPT-5 demonstrates the most sustained resilience, maintaining a 66.40% robustness score, outperforming Gemini-3-Pro by 5.59%, while other models lag behind. Data and code will be released at https://github.com/JiaQiSJTU/EvolIF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。