提出新指标量化RAG中外部知识使用程度,发现小模型在精准提取上更优。
Quantifying Prior Dominance in RAG Systems
- 用连续概率设计新指标NCU,区分真实引用与记忆回溯
- 1.5B至72B模型测试显示小模型在事实提取上表现更佳
- 大模型易受先验影响,商业API近半次冲突中忽略外部证据
检索增强生成(RAG)将大语言模型锚定于外部知识,但现有评估依赖离散启发式方法,存在‘认知盲区’——无法区分真正的上下文信息提取与参数化记忆回溯。为此,我们提出归一化上下文利用率(NCU)指标,通过零样本、理想和对抗性条件下的连续标记对数概率,严格量化上下文信息增益。在1.5B至72B参数规模的架构及一个专有商用API上进行评估发现,在严格事实提取任务中(无思维链推理),传统缩放规律表现出极端递减回报:高效的小语言模型(SLMs)表现匹配甚至超越高容量模型。此外,我们证明‘先验主导’与模型规模及专有对齐相关。所评估的商用API在近一半对抗性冲突中覆盖了明确的外部证据,且当其参数先验被反驳时,频繁出现系统性置信崩溃(负迁移)。结果凸显了小模型在严格提取工作流中的结构性认知优势与更强的上下文依从性。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。