arXiv:2507.21287cs.AI2025-07

通过多维评估提升检索增强模型的可靠性,减少幻觉。

Structured Relevance Assessment for Robust Retrieval-Augmented Language Models

  • 构建多维度评分系统,融合语义匹配与来源可信度
  • 幻觉率显著降低,推理过程更透明
  • 适合需要高可靠性的问答系统开发者

检索增强语言模型(RALMs)在减少事实性错误方面面临挑战,尤其体现在文档相关性评估和知识整合上。本文提出一种结构化相关性评估框架,通过改进文档评估、平衡内在与外部知识融合,以及有效处理无法回答的问题,提升RALM的鲁棒性。该方法采用多维度评分体系,结合嵌入式相关性评分与混合质量文档的合成训练数据,实现语义匹配与源可靠性双重考量。我们在小众主题上实施专用基准测试,设计知识整合机制,并引入对知识覆盖不足查询的“未知”响应协议。初步评估显示,幻觉率显著下降,推理过程透明度提升。该框架推动了可在动态环境中应对可变数据质量的更可靠问答系统发展。尽管在准确区分可信信息及系统延迟与全面性之间仍存挑战,本工作为提升RALM可靠性迈出重要一步。

原文摘要 · Abstract (English)

Retrieval-Augmented Language Models (RALMs) face significant challenges in reducing factual errors, particularly in document relevance evaluation and knowledge integration. We introduce a framework for structured relevance assessment that enhances RALM robustness through improved document evaluation, balanced intrinsic and external knowledge integration, and effective handling of unanswerable queries. Our approach employs a multi-dimensional scoring system that considers both semantic matching and source reliability, utilizing embedding-based relevance scoring and synthetic training data with mixed-quality documents. We implement specialized benchmarking on niche topics, a knowledge integration mechanism, and an "unknown" response protocol for queries with insufficient knowledge coverage. Preliminary evaluations demonstrate significant reductions in hallucination rates and improved transparency in reasoning processes. Our framework advances the development of more reliable question-answering systems capable of operating effectively in dynamic environments with variable data quality. While challenges persist in accurately distinguishing credible information and balancing system latency with thoroughness, this work represents a meaningful step toward enhancing RALM reliability.

检索增强幻觉抑制问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。