arXiv:2509.06631cs.CL2025-09

研究引导解码如何提升RAG系统的输出规范性与可靠性。

Guided Decoding and Its Critical Role in Retrieval-Augmented Generation

  • 对比三种引导解码方法在多轮提示下的表现
  • 发现多轮交互显著影响解码效果,存在意外性能波动
  • 为不同场景选择合适方法提供实证依据

大型语言模型(LLMs)在各类应用中的集成推动了对结构化、可靠输出的需求。检索增强生成(RAG)系统面临的核心挑战是如何使输出符合预期格式并减少幻觉。本研究考察了引导解码在RAG系统中的作用,比较了三种方法——Outlines、XGrammar和LM Format Enforcer——在不同多轮提示设置(0轮、1轮、2轮)下的表现。通过评估成功率、幻觉率和输出质量,揭示了多轮交互对引导解码的影响,发现了出人意料的性能变化,为特定应用场景的方法选择提供了依据。这项工作深化了对RAG系统中结构化输出生成的理解,兼具理论价值与实际指导意义。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into various applications has driven the need for structured and reliable responses. A key challenge in Retrieval-Augmented Generation (RAG) systems is ensuring that outputs align with expected formats while minimizing hallucinations. This study examines the role of guided decoding in RAG systems, comparing three methods, Outlines, XGrammar, and LM Format Enforcer, across different multi-turn prompting setups (0-turn, 1-turn, and 2-turn). By evaluating success rates, hallucination rates, and output quality, we provide insights into their performance and applicability. Our findings reveal how multi-turn interactions influence guided decoding, uncovering unexpected performance variations that can inform method selection for specific use cases. This work advances the understanding of structured output generation in RAG systems, offering both theoretical insights and practical guidance for LLM deployment.

RAG引导解码输出规整幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。