arXiv:2603.04370cs.AIcs.CL2026-03被引 22

测试对话智能体在复杂金融知识库中的真实交互能力

$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

  • 构建包含700篇互联文档的金融客服场景,模拟真实知识调用
  • 前沿模型在长对话中仅25.5%任务成功,且性能随尝试次数下降
  • 适合研究知识融合、推理可靠性的AI系统开发者

对话智能体在知识密集型场景中日益普及,其正确行为依赖于从大规模、专有且非结构化语料中实时检索并应用领域知识。然而现有基准多独立评估检索或工具使用,难以真实反映长周期交互中智能体的综合表现。本文提出τ-Knowledge,作为τ-Bench的扩展,用于评估智能体在需协调外部自然语言知识与工具输出以实现可验证、合规状态变更的环境中的表现。新引入的τ-Banking领域模拟真实金融科技客服流程,要求智能体在约700个互连知识文档中导航并执行工具驱动的账户更新。在基于嵌入的检索和终端搜索两种方式下,即使具备高推理预算的前沿模型,任务成功率也仅约25.5%(pass^1),且可靠性随重复试验显著下降。智能体在密集互连的知识库中难以准确检索文档,并对复杂内部策略推理错误频发。整体而言,τ-Knowledge为开发可集成非结构化知识的人机部署智能体提供了现实测试平台。

原文摘要 · Abstract (English)

Conversational agents are increasingly deployed in knowledge-intensive settings, where correct behavior depends on retrieving and applying domain-specific knowledge from large, proprietary, and unstructured corpora during live interactions with users. Yet most existing benchmarks evaluate retrieval or tool use independently of each other, creating a gap in realistic, fully agentic evaluation over unstructured data in long-horizon interactions. We introduce $τ$-Knowledge, an extension of $τ$-Bench for evaluating agents in environments where success depends on coordinating external, natural-language knowledge with tool outputs to produce verifiable, policy-compliant state changes. Our new domain, $τ$-Banking, models realistic fintech customer support workflows in which agents must navigate roughly 700 interconnected knowledge documents while executing tool-mediated account updates. Across embedding-based retrieval and terminal-based search, even frontier models with high reasoning budgets achieve only $\sim$25.5% pass^1, with reliability degrading sharply over repeated trials. Agents struggle to retrieve the correct documents from densely interlinked knowledge bases and to reason accurately over complex internal policies. Overall, $τ$-Knowledge provides a realistic testbed for developing agents that integrate unstructured knowledge in human-facing deployments.

对话系统知识检索金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。