测试大模型在长对话中持续遵守用户指令的能力
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

- 构建模拟真人对话的多轮约束测试集
- 对话越长,准确率下降超11%,多约束下降幅超40%
- 适合评估智能助手在复杂交互中的可靠性
在对话中,助手需可靠地遵循用户指令,即使指令被反复修改或矛盾。然而,现有指令遵循基准大多聚焦单轮或短多轮场景,难以评估模型在长时程对话中的表现。为此,我们提出SEQUOR,一个自动化基准,用于评估长多轮对话中约束遵循能力。SEQUOR基于真实对话提取的约束,构建模拟人格驱动的互动。结果表明,仅遵循单一约束时,随着对话长度增加,指令遵循准确率持续下降,降幅超过11%;当需同时遵守多个约束时,准确率下降超过40%;在任意时刻新增或替换约束的场景下,准确率下降超过9%。综合来看,当前模型在多轮对话中仍难以有效遵循用户指令,为评估助手的指令遵循能力提供了新方法。
原文摘要 · Abstract (English)
In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most instruction-following benchmarks focus on single-turn or short multi-turn scenarios, leaving open how well models handle long-horizon instruction-following tasks. To bridge this gap, we present SEQUOR, an automatic benchmark for evaluating constraint adherence in long multi-turn conversations. SEQUOR consists of simulated persona-driven interactions built with constraints extracted from real-world conversations. Our results show that even when following a single constraint, instruction-following accuracy consistently decreases as the conversation grows longer, with drops exceeding 11%. This decline becomes larger when models have to follow multiple constraints simultaneously, reducing their accuracy by over 40%. In scenarios where constraints are added or replaced at arbitrary points of the conversation, model accuracy decreases by more than 9%. Taken together, our results reveal that current models still struggle to follow user instructions in multi-turn conversations, and provide a way for better measuring instruction-following capabilities in assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。