首个诊断大模型代理在输入错误下的协作失效问题的多轮交互基准
Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction
- 构建多轮澄清机制,模拟真实用户交互中的模糊输入
- 实验显示输入故障导致性能显著下降,澄清效果因用户角色和故障类型而异
- 适合关注智能体安全与人机协作的研究者和开发者
随着大语言模型向自主智能体演进,用户输入常违背合作假设(如隐含意图缺失、参数不全、错误预设或表达模糊),带来执行风险,而传统文本评估无法捕捉此类风险。现有基准多假设指令明确,或仅限单轮文本澄清,难以衡量在实际任务场景中多轮消歧的能力。本文提出 extbf{Drift-Bench},首个通过多轮澄清评估智能体在状态导向与服务导向环境下的协作失效问题的诊断基准。基于经典沟通理论,该基准提供统一的合作失效分类体系,并采用角色驱动的用户模拟器与 extbf{Rise} 评估协议。实验表明,输入故障导致性能显著下降,澄清有效性随用户角色与故障类型变化。 extbf{Drift-Bench} 桥接了澄清研究与智能体安全评估,支持对可能导致不安全执行的失败进行系统性诊断。
原文摘要 · Abstract (English)
As Large Language Models transition to autonomous agents, user inputs frequently violate cooperative assumptions (e.g., implicit intent, missing parameters, false presuppositions, or ambiguous expressions), creating execution risks that text-only evaluations do not capture. Existing benchmarks typically assume well-specified instructions or restrict evaluation to text-only, single-turn clarification, and thus do not measure multi-turn disambiguation under grounded execution risk. We introduce \textbf{Drift-Bench}, the first diagnostic benchmark that evaluates agentic pragmatics under input faults through multi-turn clarification across state-oriented and service-oriented execution environments. Grounded in classical theories of communication, \textbf{Drift-Bench} provides a unified taxonomy of cooperative breakdowns and employs a persona-driven user simulator with the \textbf{Rise} evaluation protocol. Experiments show substantial performance drops under these faults, with clarification effectiveness varying across user personas and fault types. \MethodName bridges clarification research and agent safety evaluation, enabling systematic diagnosis of failures that can lead to unsafe executions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。