arXiv:2604.04269cs.AIcs.LG2026-04被引 2

提出可靠智能体检索需关注过程正确性,而非仅语言流畅度。

Beyond Fluency: Toward Reliable Trajectories in Agentic IR

  • 在多步推理-行动-观察循环中设置验证关卡,阻断错误传播。
  • 早期小错误会级联导致结果偏差,即使输出仍流畅。
  • 适合构建安全可控的智能体系统,尤其工业级应用者。

信息检索正从被动文档排序转向自主智能体工作流,其在多步推理-行动-观察循环中运行。在此长周期轨迹中,早期微小错误可能级联,导致内部推理与外部工具执行出现功能错位,尽管语言表达持续流畅。本文总结工业级智能体系统中观察到的失效模式,将错误归类于规划、检索、推理和执行阶段。我们主张,安全部署需超越终点准确率,转向轨迹完整性与因果归因。为应对误差累积与虚假流畅性,提出在每个交互单元设置验证关卡,并倡导在可校准的不确定性下系统性拒绝执行。可靠的智能体信息检索系统必须优先保证过程正确性和基于事实的执行,而非看似合理却未经验证的完成。

原文摘要 · Abstract (English)

Information Retrieval is shifting from passive document ranking toward autonomous agentic workflows that operate in multi-step Reason-Act-Observe loops. In such long-horizon trajectories, minor early errors can cascade, leading to functional misalignment between internal reasoning and external tool execution despite continued linguistic fluency. This position paper synthesizes failure modes observed in industrial agentic systems, categorizing errors across planning, retrieval, reasoning, and execution. We argue that safe deployment requires moving beyond endpoint accuracy toward trajectory integrity and causal attribution. To address compounding error and deceptive fluency, we propose verification gates at each interaction unit and advocate systematic abstention under calibrated uncertainty. Reliable Agentic IR systems must prioritize process correctness and grounded execution over plausible but unverified completion.

智能体信息检索可靠性验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。