用智能体与语义缓存减少大模型幻觉,关键在明确上下文。
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

- 通过嵌套学习与语义缓存构建三阶段检测流程。
- 上下文明确性提升贡献83.5%的幻觉改善,事实密度未变。
- 仅第一阶段有幻觉标记,适合关注可解释性的研究者。
本文提出一种基于HOPE启发的嵌套学习架构,结合连续记忆系统(CMS)与语义相似性缓存,用于幻觉检测与缓解。在包含310个提示的混合基准上测试(217个知识不确定性提示,93个虚构诱导压力测试),采用开放楼层协议(Open Floor Protocol)协调三阶段流程,评估五个关键绩效指标。总幻觉得分整体提升6.1%的可实现范围,其中83.5%的改进源于“显式上下文化”维度;事实陈述密度(最接近无支持内容的维度)保持不变。97.7%的改进发生在第一轮审查阶段。第五项指标“可观测性”单独报告:该指标反映是否存在OFP注释通道,在审查阶段上升147%,但在最终阶段回落,因该阶段不传播注释通道。三位标注员独立标注93个压力测试的最终响应,10例(10.8%,95%置信区间5.9-18.7)仍将虚构内容当作真实,显著性水平α=0.586,低于常规阈值;显式上下文化与标注结果单调一致。使用Llama 3.1、Gemma 4和Qwen 3作为裁判重新评分,验证了改进效果,并显示跨家族裁判优于原评估器(斯皮尔曼相关系数rho=-0.772 vs -0.477)。语义缓存服务于47.7%的模型调用。总体而言,名义上的多维可靠性评分实为单维驱动,且仅此维度获得外部支持。
原文摘要 · Abstract (English)
This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuum Memory Systems (CMS) and semantic similarity caching, tested on a hybrid benchmark of 310 prompts (217 epistemic-uncertainty prompts, 93 fabrication-induction stress tests). A three-stage pipeline orchestrated via the Open Floor Protocol is evaluated with five KPIs; four score the response and aggregate into a Total Hallucination Score. The score improves end to end by 6.1% of its attainable range, but 83.5% of that gain is attributable to a single dimension, Explicit Contextualization, while Factual Claim Density, the dimension closest to unsupported content, stays flat; 97.7% of the gain arrives at the first review stage. The fifth indicator, observability, is reported separately, since it registers the presence of an OFP annotation channel rather than a property of the response: it rises 147% at the review stage, the only stage carrying explicit hallucination markers, then falls back at the final stage, which does not propagate the channel. Three annotators independently labelled every final-stage response on the 93 stress prompts: in 10 of 93 cases (10.8%, 95% CI 5.9-18.7) the final answer still presents the invented item as real, at alpha=0.586, below the conventional threshold; Explicit Contextualization tracks these labels monotonically. Re-scoring all 930 outputs with Llama 3.1, Gemma 4, and Qwen 3 as judges confirms the gain and ranks the cross-family judges above the original evaluator against human labels (rho=-0.772 vs -0.477). Semantic caching serves 47.7% of model calls. A nominally multi-dimensional reliability score is thus effectively one-dimensional, and only that dimension has external support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。